Wikitech
labswiki
https://wikitech.wikimedia.org/wiki/Main_Page
MediaWiki 1.47.0-wmf.19
first-letter
Media
Special
Talk
User
User talk
Wikitech
Wikitech talk
File
File talk
MediaWiki
MediaWiki talk
Template
Template talk
Help
Help talk
Category
Category talk
Obsolete
Obsolete talk
OfficeIT
OfficeIT talk
Tool
Tool talk
Nova Resource
Nova Resource Talk
Heira
Heira Talk
TimedText
TimedText talk
Module
Module talk
PoolCounter
0
4105
2456806
1801848
2026-09-11T20:35:23Z
RLazarus (WMF)
15215
Redirected page to [[SRE/Service Operations/Documentation/Reboots#PoolCounter]]
2456806
wikitext
text/x-wiki
#REDIRECT [[SRE/Service_Operations/Documentation/Reboots#PoolCounter]]
ltx5pl9fsw6528790jqyqbjqmghizc6
Poolcounter
0
4107
2456807
48739
2026-09-11T20:36:01Z
RLazarus (WMF)
15215
Changed redirect target from [[PoolCounter]] to [[SRE/Service Operations/Documentation/Reboots#PoolCounter]]
2456807
wikitext
text/x-wiki
#REDIRECT [[SRE/Service_Operations/Documentation/Reboots#PoolCounter]]
ltx5pl9fsw6528790jqyqbjqmghizc6
Deployments
0
4108
2456811
2456690
2026-09-12T02:00:17Z
DeploymentCalendarTool
20896
Remove Week of September 07
2456811
wikitext
text/x-wiki
{{Navigation MediaWiki deployment}}
This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]].
== Getting started ==
Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there.
If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>).
* '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule.
* '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]].
* '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join.
** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div>
* Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks.
**To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>.
**To create an one-off window, simply edit this page accordingly
** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies.
** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]].
* '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority.
__TOC__
{{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}}
[[Category:Deployment]]
{{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}}
== Week of September 21 ==
=== {{Deployment_day|date=2026-09-23}} ===
{{Deployment calendar event card
|when=2026-09-23 13:00 UTC
|length=4
|window=Datacenter switch over
|who={{ircnick|slyngs}}
|what=Datacenter switch over
}}
==Week of September 14==
==={{Deployment_day|date=2026-09-13}}===
{{Deployment calendar event card
|when=2026-09-13 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==={{Deployment_day|date=2026-09-14}}===
{{Deployment calendar event card
|when=2026-09-14 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-14 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-14 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-14 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|Hide_on_rosie|Hide_on_rosie}}
{{deploy|type=config|gerrit=1338902|title=afwiki: Create Draft and Draft talk namespaces|status=}} - {{phabricator|T437576}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-14 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-14 08:30 SF
|length=0.5
|window=Wikimedia Portals Update
|who={{ircnick|jan_drewniak|Jan Drewniak}}
|what=Weekly window for the portals page: https://www.wikipedia.org/
}}
{{Deployment calendar event card
|when=2026-09-14 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-14 10:00 SF
|length=0.5
|window=Wikidata Query Service weekly deploy
|who={{ircnick|ryankemper|Ryan}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-14 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-14 14:00 SF
|length=2
|window=Weekly Security deployment window
|who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}}
|what=Held deployment window for Security-team related deploys.
}}
{{Deployment calendar event card
|when=2026-09-14 16:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-14 19:00 SF
|length=1
|window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Branch <code>wmf/1.47.0-wmf.20</code>
}}
{{Deployment calendar event card
|when=2026-09-14 20:00 SF
|length=1
|window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Deploy <code>wmf/1.47.0-wmf.20</code> to testwikis
}}
{{Deployment calendar event card
|when=2026-09-14 21:00 SF
|length=1
|window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version)
|who=N/A
|what=Runs <code>scap clean auto</code>
}}
{{Deployment calendar event card
|when=2026-09-14 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-14 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-15}}===
{{Deployment calendar event card
|when=2026-09-15 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-15 01:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19|1.47.0-wmf.19}}
* group0 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-15 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-15 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-15 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-15 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-15 07:00 SF
|length=0.5
|window=Test Kitchen UI Deployment Window
|who=Experimentation Platform Team
|what=Deployment of Test Kitchen UI (fka MPIC)
}}
{{Deployment calendar event card
|when=2026-09-15 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-15 08:00 SF
|length=1
|window=SRE Collaboration Services office hours
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=Services including Gerrit, Phorge (Phabricator), GitLab
}}
{{Deployment calendar event card
|when=2026-09-15 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-15 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-15 11:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot)
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19|1.47.0-wmf.19}}
* group0 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-15 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-15 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-15 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-16}}===
{{Deployment calendar event card
|when=2026-09-16 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-16 01:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19}}
* group1 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-16 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-16 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-16 04:00 SF
|length=1
|window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]]
|who=Marielle ({{ircnick|mvolz}})
|what=See [[mw:Citoid|Citoid]]
}}
{{Deployment calendar event card
|when=2026-09-16 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-16 07:00 SF
|length=1
|window=Wikifunctions Services UTC Afternoon
|who=Abstract Wikipedia team (Africa, Europe, Eastern Americas)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-16 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-16 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-16 11:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot)
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19}}
* group1 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-16 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-16 14:00 SF
|length=1
|window=Wikifunctions Services UTC Late
|who=Abstract Wikipedia team (North and South America)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-16 15:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-16 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-16 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-17}}===
{{Deployment calendar event card
|when=2026-09-17 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-17 01:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20}}
* group2 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-17 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-17 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-17 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-17 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-17 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-17 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-17 10:00 SF
|length=1
|window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker)
|who={{ircnick|bd808}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-17 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-17 11:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot)
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20}}
* group2 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-17 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-17 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-17 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-18}}===
{{Deployment calendar event card
|when=2026-09-18 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
{{Deployment calendar event card
|when=2026-09-18 04:00 SF
|length=0.5
|window=GitLab version upgrades
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=GitLab version upgrades
}}
==={{Deployment_day|date=2026-09-19}}===
{{Deployment calendar event card
|when=2026-09-19 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==Week of September 21==
==={{Deployment_day|date=2026-09-20}}===
{{Deployment calendar event card
|when=2026-09-20 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==={{Deployment_day|date=2026-09-21}}===
{{Deployment calendar event card
|when=2026-09-21 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-21 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-21 08:30 SF
|length=0.5
|window=Wikimedia Portals Update
|who={{ircnick|jan_drewniak|Jan Drewniak}}
|what=Weekly window for the portals page: https://www.wikipedia.org/
}}
{{Deployment calendar event card
|when=2026-09-21 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 10:00 SF
|length=0.5
|window=Wikidata Query Service weekly deploy
|who={{ircnick|ryankemper|Ryan}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-21 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 14:00 SF
|length=2
|window=Weekly Security deployment window
|who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}}
|what=Held deployment window for Security-team related deploys.
}}
{{Deployment calendar event card
|when=2026-09-21 16:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-21 19:00 SF
|length=1
|window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Branch <code>wmf/0.00.0-wmf.0</code>
}}
{{Deployment calendar event card
|when=2026-09-21 20:00 SF
|length=1
|window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Deploy <code>wmf/0.00.0-wmf.0</code> to testwikis
}}
{{Deployment calendar event card
|when=2026-09-21 21:00 SF
|length=1
|window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version)
|who=N/A
|what=Runs <code>scap clean auto</code>
}}
{{Deployment calendar event card
|when=2026-09-21 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-22}}===
{{Deployment calendar event card
|when=2026-09-22 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-22 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-22 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-22 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 07:00 SF
|length=0.5
|window=Test Kitchen UI Deployment Window
|who=Experimentation Platform Team
|what=Deployment of Test Kitchen UI (fka MPIC)
}}
{{Deployment calendar event card
|when=2026-09-22 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-22 08:00 SF
|length=1
|window=SRE Collaboration Services office hours
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=Services including Gerrit, Phorge (Phabricator), GitLab
}}
{{Deployment calendar event card
|when=2026-09-22 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-22 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-22 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0->0.00.0-wmf.0|0.00.0-wmf.0|0.00.0-wmf.0}}
* group0 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-22 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-22 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-23}}===
{{Deployment calendar event card
|when=2026-09-23 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-23 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 04:00 SF
|length=1
|window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]]
|who=Marielle ({{ircnick|mvolz}})
|what=See [[mw:Citoid|Citoid]]
}}
{{Deployment calendar event card
|when=2026-09-23 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 07:00 SF
|length=1
|window=Wikifunctions Services UTC Afternoon
|who=Abstract Wikipedia team (Africa, Europe, Eastern Americas)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-23 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-23 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0|0.00.0-wmf.0->0.00.0-wmf.0|0.00.0-wmf.0}}
* group1 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-23 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 14:00 SF
|length=1
|window=Wikifunctions Services UTC Late
|who=Abstract Wikipedia team (North and South America)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-23 15:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-23 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-24}}===
{{Deployment calendar event card
|when=2026-09-24 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-24 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-24 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-24 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-24 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-24 10:00 SF
|length=1
|window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker)
|who={{ircnick|bd808}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-24 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-24 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0|0.00.0-wmf.0|0.00.0-wmf.0->0.00.0-wmf.0}}
* group2 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-24 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-24 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-25}}===
{{Deployment calendar event card
|when=2026-09-25 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
{{Deployment calendar event card
|when=2026-09-25 04:00 SF
|length=0.5
|window=GitLab version upgrades
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=GitLab version upgrades
}}
==={{Deployment_day|date=2026-09-26}}===
{{Deployment calendar event card
|when=2026-09-26 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
k5xchgnfrvrfd4jf5h1oymlk821j9gt
2456815
2456811
2026-09-12T02:56:55Z
ScheduleDeploymentBot
37566
Add [[gerrit:1340213]] to Monday, September 14 UTC afternoon backport window
2456815
wikitext
text/x-wiki
{{Navigation MediaWiki deployment}}
This page tracks '''upcoming''' '''deployments''' of software to the [[:m:Special:SiteMatrix|Wikimedia Foundation servers]].
== Getting started ==
Ensure you joined the {{irc|wikimedia-operations}} IRC channel as all deployment-related communications happen there.
If you need help, contact [[:mw:Wikimedia Release Engineering Team|Release Engineering]] on IRC at {{irc|wikimedia-releng}}; and ping Tyler (<code>thcipriani</code>).
* '''MediaWiki is deployed weekly''' through the [[/Train|Deployment Train]]. Other services follow their own schedule.
* '''Times are pinned to San Francisco''', thus the UTC time changes in March and November per [[:en:Daylight saving time in the United States|DST]].
* '''Prefer regular [[Backport windows]]''' over adding new windows. To request deployment of a config change or backport, add your username and Gerrit URL to one of the backport windows on this page. You must be online in #wikimedia-operations on IRC during your deployment and install [[WikimediaDebug]] ahead of time. The #wikimedia-operations channel requires you to [[:m:IRC/Instructions#Register your nickname, identify, and enforce|register your nickname]] before you can join.
** You can use the '''backport scheduling tool''' to more easily edit this page: <div style="text-align: center; margin: 1em 0">{{Clickable button 2|:toollabs:schedule-deployment|Schedule a backport|class=mw-ui-progressive}}</div>
* Tasks that meet [[/Inclusion criteria|Inclusion criteria]] '''require their own windows''', which includes long-running tasks. '''Schedule more time''' than you think you need to account for delays and set backs, we recommend one hour for most tasks.
**To create or modify a recurring deploy window, send a patchset to [[:gitlab:repos/releng/release/-/blob/main/make-deployment-calendar/deployments-calendar.yaml|deployments-calendar.yaml file]] in <code>repos/releng/release.git</code>.
**To create an one-off window, simply edit this page accordingly
** '''Announce''' changes to the [[mail:ops|ops mailing list]] ahead of time if you anticipate or are uncertain about noticeable impacts to database load, HTTP caching, or the introduction of new cookies.
** '''Announce''' deployments of major features to the community via [[:m:Tech/News/Next|Tech News]] and/or via other [[:mw:Wikimedia_Product_Guidance/Communication_channels|Product communication channels]].
* '''Something went wrong?''' See [[Incident response]]. Is there a user-impacting problem? Communicate in the {{irc|wikimedia-operations}} IRC channel. If there is a Phabricator task, ensure [[:phab:tag/wikimedia-incident/|#Wikimedia-Incident]] is tagged, and consider setting the [[:mw:Phabricator/Project_management#Priority_levels|Unbreak Now]] priority.
__TOC__
{{anchor|Next Week|Near Term|Near term|Near-term}}{{clear}}
[[Category:Deployment]]
{{Note|content=Subscribe in Google Calendar via <code>wikimedia.org_rudis09ii2mm5fk4hgdjeh1u64@group.calendar.google.com</code>.<br>This may not include one-off windows. '''If there are differences, then the wiki page is canonical and correct'''.}}
== Week of September 21 ==
=== {{Deployment_day|date=2026-09-23}} ===
{{Deployment calendar event card
|when=2026-09-23 13:00 UTC
|length=4
|window=Datacenter switch over
|who={{ircnick|slyngs}}
|what=Datacenter switch over
}}
==Week of September 14==
==={{Deployment_day|date=2026-09-13}}===
{{Deployment calendar event card
|when=2026-09-13 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==={{Deployment_day|date=2026-09-14}}===
{{Deployment calendar event card
|when=2026-09-14 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-14 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-14 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-14 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|Hide_on_rosie|Hide_on_rosie}}
{{deploy|type=config|gerrit=1338902|title=afwiki: Create Draft and Draft talk namespaces|status=}} - {{phabricator|T437576}}
{{ircnick|Daimona|Daimona}}
{{deploy|type=config|gerrit=1340213|title=Enable $wgCampaignEventsEnableWorklistEventDiscoveryTracking in beta|status=}} - {{phabricator|T434512}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-14 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-14 08:30 SF
|length=0.5
|window=Wikimedia Portals Update
|who={{ircnick|jan_drewniak|Jan Drewniak}}
|what=Weekly window for the portals page: https://www.wikipedia.org/
}}
{{Deployment calendar event card
|when=2026-09-14 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-14 10:00 SF
|length=0.5
|window=Wikidata Query Service weekly deploy
|who={{ircnick|ryankemper|Ryan}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-14 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-14 14:00 SF
|length=2
|window=Weekly Security deployment window
|who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}}
|what=Held deployment window for Security-team related deploys.
}}
{{Deployment calendar event card
|when=2026-09-14 16:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-14 19:00 SF
|length=1
|window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Branch <code>wmf/1.47.0-wmf.20</code>
}}
{{Deployment calendar event card
|when=2026-09-14 20:00 SF
|length=1
|window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Deploy <code>wmf/1.47.0-wmf.20</code> to testwikis
}}
{{Deployment calendar event card
|when=2026-09-14 21:00 SF
|length=1
|window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version)
|who=N/A
|what=Runs <code>scap clean auto</code>
}}
{{Deployment calendar event card
|when=2026-09-14 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-14 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-15}}===
{{Deployment calendar event card
|when=2026-09-15 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-15 01:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19|1.47.0-wmf.19}}
* group0 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-15 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-15 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-15 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-15 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-15 07:00 SF
|length=0.5
|window=Test Kitchen UI Deployment Window
|who=Experimentation Platform Team
|what=Deployment of Test Kitchen UI (fka MPIC)
}}
{{Deployment calendar event card
|when=2026-09-15 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-15 08:00 SF
|length=1
|window=SRE Collaboration Services office hours
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=Services including Gerrit, Phorge (Phabricator), GitLab
}}
{{Deployment calendar event card
|when=2026-09-15 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-15 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-15 11:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot)
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19|1.47.0-wmf.19}}
* group0 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-15 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-15 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-15 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-16}}===
{{Deployment calendar event card
|when=2026-09-16 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-16 01:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19}}
* group1 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-16 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-16 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-16 04:00 SF
|length=1
|window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]]
|who=Marielle ({{ircnick|mvolz}})
|what=See [[mw:Citoid|Citoid]]
}}
{{Deployment calendar event card
|when=2026-09-16 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-16 07:00 SF
|length=1
|window=Wikifunctions Services UTC Afternoon
|who=Abstract Wikipedia team (Africa, Europe, Eastern Americas)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-16 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-16 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-16 11:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot)
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20|1.47.0-wmf.19}}
* group1 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-16 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-16 14:00 SF
|length=1
|window=Wikifunctions Services UTC Late
|who=Abstract Wikipedia team (North and South America)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-16 15:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-16 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-16 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-17}}===
{{Deployment calendar event card
|when=2026-09-17 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-17 01:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20}}
* group2 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-17 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-17 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-17 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-17 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-17 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-17 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-17 10:00 SF
|length=1
|window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker)
|who={{ircnick|bd808}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-17 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-17 11:00 SF
|length=2
|window=MediaWiki train - Utc-0+Utc-7 Version (secondary timeslot)
|who={{ircnick|jnuche|Jaime}}, {{ircnick|dduvall|Dan}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.20|1.47.0-wmf.20|1.47.0-wmf.19->1.47.0-wmf.20}}
* group2 to [[mw:MediaWiki_1.47/wmf.20|1.47.0-wmf.20]]
* '''Blockers: {{phabricator|T430839}}'''
}}
{{Deployment calendar event card
|when=2026-09-17 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-17 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-17 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-18}}===
{{Deployment calendar event card
|when=2026-09-18 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
{{Deployment calendar event card
|when=2026-09-18 04:00 SF
|length=0.5
|window=GitLab version upgrades
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=GitLab version upgrades
}}
==={{Deployment_day|date=2026-09-19}}===
{{Deployment calendar event card
|when=2026-09-19 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==Week of September 21==
==={{Deployment_day|date=2026-09-20}}===
{{Deployment calendar event card
|when=2026-09-20 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==={{Deployment_day|date=2026-09-21}}===
{{Deployment calendar event card
|when=2026-09-21 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-21 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-21 08:30 SF
|length=0.5
|window=Wikimedia Portals Update
|who={{ircnick|jan_drewniak|Jan Drewniak}}
|what=Weekly window for the portals page: https://www.wikipedia.org/
}}
{{Deployment calendar event card
|when=2026-09-21 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 10:00 SF
|length=0.5
|window=Wikidata Query Service weekly deploy
|who={{ircnick|ryankemper|Ryan}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-21 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-21 14:00 SF
|length=2
|window=Weekly Security deployment window
|who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}}
|what=Held deployment window for Security-team related deploys.
}}
{{Deployment calendar event card
|when=2026-09-21 16:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-21 19:00 SF
|length=1
|window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Branch <code>wmf/0.00.0-wmf.0</code>
}}
{{Deployment calendar event card
|when=2026-09-21 20:00 SF
|length=1
|window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Deploy <code>wmf/0.00.0-wmf.0</code> to testwikis
}}
{{Deployment calendar event card
|when=2026-09-21 21:00 SF
|length=1
|window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version)
|who=N/A
|what=Runs <code>scap clean auto</code>
}}
{{Deployment calendar event card
|when=2026-09-21 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-21 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-22}}===
{{Deployment calendar event card
|when=2026-09-22 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-22 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-22 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-22 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 07:00 SF
|length=0.5
|window=Test Kitchen UI Deployment Window
|who=Experimentation Platform Team
|what=Deployment of Test Kitchen UI (fka MPIC)
}}
{{Deployment calendar event card
|when=2026-09-22 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-22 08:00 SF
|length=1
|window=SRE Collaboration Services office hours
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=Services including Gerrit, Phorge (Phabricator), GitLab
}}
{{Deployment calendar event card
|when=2026-09-22 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-22 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-22 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0->0.00.0-wmf.0|0.00.0-wmf.0|0.00.0-wmf.0}}
* group0 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-22 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-22 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-22 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-23}}===
{{Deployment calendar event card
|when=2026-09-23 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-23 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 04:00 SF
|length=1
|window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]]
|who=Marielle ({{ircnick|mvolz}})
|what=See [[mw:Citoid|Citoid]]
}}
{{Deployment calendar event card
|when=2026-09-23 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 07:00 SF
|length=1
|window=Wikifunctions Services UTC Afternoon
|who=Abstract Wikipedia team (Africa, Europe, Eastern Americas)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-23 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-23 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0|0.00.0-wmf.0->0.00.0-wmf.0|0.00.0-wmf.0}}
* group1 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-23 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-23 14:00 SF
|length=1
|window=Wikifunctions Services UTC Late
|who=Abstract Wikipedia team (North and South America)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-23 15:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-23 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-23 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-24}}===
{{Deployment calendar event card
|when=2026-09-24 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 02:00 UTC
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-24 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-24 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-24 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-24 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-24 10:00 SF
|length=1
|window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker)
|who={{ircnick|bd808}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-24 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-24 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|thcipriani|Tyler}}, {{ircnick|thcipriani|Tyler}}
|what=[[mw:MediaWiki 1.00/Roadmap#Schedule for the deployments|1.00 schedule]]
{{DeployOneWeekMini|0.00.0-wmf.0|0.00.0-wmf.0|0.00.0-wmf.0->0.00.0-wmf.0}}
* group2 to [[mw:MediaWiki_1.00/wmf.0|0.00.0-wmf.0]]
* '''Blockers: {{phabricator|None}}'''
}}
{{Deployment calendar event card
|when=2026-09-24 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-24 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-24 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-25}}===
{{Deployment calendar event card
|when=2026-09-25 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
{{Deployment calendar event card
|when=2026-09-25 04:00 SF
|length=0.5
|window=GitLab version upgrades
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=GitLab version upgrades
}}
==={{Deployment_day|date=2026-09-26}}===
{{Deployment calendar event card
|when=2026-09-26 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
1cwjdjuafxj2qtoxgyevao5w6ovsjkd
Server Admin Log
0
7919
2456767
2456764
2026-09-11T13:10:45Z
Stashbot
7414
brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
2456767
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
gi3awlo9p12ee17x174pmrfyt7gafw5
2456768
2456767
2026-09-11T13:11:20Z
Stashbot
7414
brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
2456768
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
0k5a606x49s70fxl2z7a24ohpvjih4h
2456769
2456768
2026-09-11T13:15:48Z
Stashbot
7414
brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
2456769
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
1d1xrdeznflfny818a0ywnvk61ob3pa
2456770
2456769
2026-09-11T13:16:43Z
Stashbot
7414
brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
2456770
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
9xtal2gi32crn57o1rmlut32i80xxfn
2456771
2456770
2026-09-11T13:28:56Z
Stashbot
7414
bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
2456771
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
qoc3mfbayldvhm4sjipkbt6ndqtmiag
2456772
2456771
2026-09-11T13:33:23Z
Stashbot
7414
sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
2456772
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
3h5ytaekt49lj61d92wyxmsfcua4qwz
2456773
2456772
2026-09-11T13:33:46Z
Stashbot
7414
sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
2456773
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
htqtrevx9a59ide9cd31qgwms7d0t69
2456774
2456773
2026-09-11T13:40:52Z
Stashbot
7414
bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
2456774
wikitext
text/x-wiki
== 2026-09-11 ==
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
l70dytqnjuipxgorj3csk08muskc1ng
2456775
2456774
2026-09-11T14:09:55Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
2456775
wikitext
text/x-wiki
== 2026-09-11 ==
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
jzl9phxsw1tfxaog6zynx9y5g365n0i
2456776
2456775
2026-09-11T14:10:23Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
2456776
wikitext
text/x-wiki
== 2026-09-11 ==
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
k05cq3ktgby46kw8enkodn1urqjtz17
2456777
2456776
2026-09-11T14:10:25Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
2456777
wikitext
text/x-wiki
== 2026-09-11 ==
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
1unvg59wk38cvrauh8t2t052twlg8d6
2456778
2456777
2026-09-11T14:10:53Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
2456778
wikitext
text/x-wiki
== 2026-09-11 ==
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
t90rrnhr5y9xpiacamk8ef994snssl9
2456779
2456778
2026-09-11T14:39:44Z
Stashbot
7414
bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
2456779
wikitext
text/x-wiki
== 2026-09-11 ==
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
g70qucjvxxgf5ccicx3pqvtd1jr4o2n
2456783
2456779
2026-09-11T16:08:01Z
Stashbot
7414
bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
2456783
wikitext
text/x-wiki
== 2026-09-11 ==
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
8e2w69vgmmm0ulugjyt9am0pguc4n3e
2456787
2456783
2026-09-11T16:40:27Z
Stashbot
7414
sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808|Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813|Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
2456787
wikitext
text/x-wiki
== 2026-09-11 ==
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
htnpfytsn489fto88geunmx2a4h8rvq
2456788
2456787
2026-09-11T16:42:43Z
Stashbot
7414
sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808|Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813|Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
2456788
wikitext
text/x-wiki
== 2026-09-11 ==
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
kvqjjhcbqbuoicz3yvd9zgoaunayix8
2456790
2456788
2026-09-11T16:43:16Z
Stashbot
7414
sbassett@deploy1003: sbassett: Continuing with deployment
2456790
wikitext
text/x-wiki
== 2026-09-11 ==
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
fxppso7bz48cruh82d6av3twqauqnx7
2456791
2456790
2026-09-11T16:47:51Z
Stashbot
7414
sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808|Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813|Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
2456791
wikitext
text/x-wiki
== 2026-09-11 ==
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
dhbg327rfgax62egk0nfk0nyhu232bb
2456808
2456791
2026-09-11T21:50:37Z
Stashbot
7414
amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
2456808
wikitext
text/x-wiki
== 2026-09-11 ==
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
d1jafwav0sjh00q7heakv2a352ixcgn
2456809
2456808
2026-09-11T21:51:56Z
Stashbot
7414
amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
2456809
wikitext
text/x-wiki
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
hvei44nit0yucdi4xoc6s7vlqondc2l
2456813
2456809
2026-09-12T02:00:57Z
Stashbot
7414
mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2456813
wikitext
text/x-wiki
== 2026-09-12 ==
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
40ihz8gj5rivpv9gq4i7tiva4ja9vvj
2456814
2456813
2026-09-12T02:08:32Z
Stashbot
7414
mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
2456814
wikitext
text/x-wiki
== 2026-09-12 ==
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 35s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-11 ==
* 21:51 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 21:50 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:47 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] (duration: 07m 23s)
* 16:43 sbassett@deploy1003: sbassett: Continuing with deployment
* 16:42 sbassett@deploy1003: sbassett: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:40 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1339808{{!}}Use escaped() for story link parentheses in recent changes (T182213)]], [[gerrit:1339813{{!}}Use escaped() for HTML parentheses params in ChangeLineFormatter (T182213)]]
* 16:08 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:39 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 14:10 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:10 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:09 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:40 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:33 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:28 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 13:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: db1199 repool
* 11:05 moritzm: installing Linux 6.1.187 on Bookworm hosts
* 11:05 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Jcrespo out of all services on: 2443 hosts
* 10:44 aokoth@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2020 in turn
* 10:43 Emperor: restart versitygw@objectstorage0[0-3].service on backup2019 in turn
* 10:41 Emperor: restart versitygw@objectstorage0[0-3].service on backup2018 in turn
* 10:40 Emperor: restart versitygw@objectstorage0[0-3].service on backup2017 in turn
* 10:39 Emperor: restart versitygw@objectstorage0[0-3].service on backup2016 in turn
* 10:37 Emperor: restart versitygw@objectstorage0[0-3].service on backup2015 in turn
* 10:36 Emperor: restart versitygw@objectstorage0[0-3].service on backup1020 in turn
* 10:35 Emperor: restart versitygw@objectstorage0[0-3].service on backup1019 in turn
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1199: db1199 repool
* 10:33 Emperor: restart versitygw@objectstorage0[0-3].service on backup1018 in turn
* 10:32 Emperor: restart versitygw@objectstorage0[0-3].service on backup1017 in turn
* 10:30 Emperor: restart versitygw@objectstorage0[0-3].service on backup1016 in turn
* 10:20 Emperor: restart versitygw@objectstorage0[1-3].service on backup1015 in turn
* 10:17 Emperor: restart versitygw@objectstorage00.service on backup1015
* 10:15 aokoth@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T437680|T437680]]
* 08:46 slyngs: Update CAS/SSO to CAS 7.3.8.3
* 08:45 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:45 slyngshede@dns1004: END - running authdns-update
* 08:43 slyngshede@dns1004: START - running authdns-update
* 08:36 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:28 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 7 hosts with reason: Restarting s5
* 08:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db[1154,1269].eqiad.wmnet with reason: Restarting s5
* 08:20 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:20 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 08:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Repooling db1159
* 07:58 arnaudb@cumin1003: END (FAIL) - Cookbook sre.gitlab.upgrade (exit_code=99) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:58 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:54 arnaudb@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:29 arnaudb@cumin1003: END (ERROR) - Cookbook sre.gitlab.upgrade (exit_code=97) on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:28 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Upgrade gitlab
* 07:27 arnaudb@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Upgrade gitlab
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1199: Needs to clone another host from this one
* 07:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1199: Needs to clone another host from this one
* 07:16 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Repooling db1159
* 07:15 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1199.eqiad.wmnet with reason: Cloning s4
* 07:10 TimStarling: killed jobs for [[phab:T437056|T437056]] since they weren't purging
* 06:38 TimStarling: also started refreshLinks for ptwiki and zhwiki, reparsing ~3000 pages altogether [[phab:T437056|T437056]]
* 06:27 TimStarling: for [[phab:T437056|T437056]]: mwscript-k8s refreshLinks.php --wiki=eswiki --tracking-category scribunto-common-error-category
* 05:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1159: Needs to clone another host from this one
* 05:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1159: Needs to clone another host from this one
* 05:30 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1159.eqiad.wmnet with reason: Cloning
* 05:29 marostegui: Start cloning db1245:s5 [[phab:T437563|T437563]]
* 05:27 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet,db1245.eqiad.wmnet with reason: Cloning
* 05:25 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] (duration: 09m 59s)
* 05:21 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:20 tstarling@deploy1003: tstarling: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:15 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1339441{{!}}Revert Lua 5.4 support patches (T437056)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 50s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-10 ==
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1349.eqiad.wmnet
* 23:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1349.eqiad.wmnet
* 23:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1349.eqiad.wmnet
* 23:09 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1349.eqiad.wmnet with reason: host reimage
* 22:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1349
* 22:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1349
* 22:32 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1349
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1349.eqiad.wmnet 198.48.64.10.in-addr.arpa 8.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1349 - rzl@cumin2003"
* 22:28 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1349
* 22:27 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1349.eqiad.wmnet with OS trixie
* 22:27 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1349.eqiad.wmnet
* 22:26 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1349.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1348.eqiad.wmnet
* 22:25 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1348.eqiad.wmnet
* 22:23 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2006.codfw.wmnet with reason: host reimage
* 22:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 22:12 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] (duration: 10m 59s)
* 22:06 musikanimal@deploy1003: kemayo, musikanimal: Rolling back deployment
* 22:05 musikanimal@deploy1003: kemayo, musikanimal: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:01 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1322175{{!}}CodeMirror: turn on the new 2017 editor integration (T432558)]], [[gerrit:1338308{{!}}[CodeMirror] enable for new and logged-out users by default on enwiki (T288161)]]
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 22:00 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:52 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1348.eqiad.wmnet with reason: host reimage
* 21:47 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2006.codfw.wmnet with OS bookworm
* 21:47 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] (duration: 13m 23s)
* 21:46 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:46 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:44 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2006.codfw.wmnet
* 21:42 derenrich@deploy1003: derenrich: Continuing with deployment
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1348
* 21:40 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1348
* 21:39 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1348
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1348.eqiad.wmnet 197.48.64.10.in-addr.arpa 7.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:39 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:39 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1348 - rzl@cumin2003"
* 21:37 derenrich@deploy1003: derenrich: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:35 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1348
* 21:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1348.eqiad.wmnet with OS trixie
* 21:33 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1348.eqiad.wmnet
* 21:33 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338439{{!}}Enable discord preview extension code on testwiki (tk 2) (T437344 T431352)]]
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1348.eqiad.wmnet
* 21:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1348.eqiad.wmnet
* 21:31 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2006.codfw.wmnet
* 21:31 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] (duration: 09m 45s)
* 21:27 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:26 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1337984{{!}}Enable ReadingLists on mediawiki and wikitech (T437113)]]
* 21:19 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2006.codfw.wmnet
* 21:19 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2006.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 21:17 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] (duration: 13m 54s)
* 21:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2005.codfw.wmnet with OS bookworm
* 21:12 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:07 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdeb
* 21:03 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1339076{{!}}Instrument donor ID dialog: hooks defined (T435565)]], [[gerrit:1339075{{!}}Instrument the donor account dialog (T435565)]], [[gerrit:1338301{{!}}DonorIdentification: Adjust experiment behavior (T435534)]], [[gerrit:1338287{{!}}Allow donor consent workflow to work on non-Minerva skins (T436572)]]
* 20:54 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:52 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] (duration: 23m 49s)
* 20:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2005.codfw.wmnet with reason: host reimage
* 20:47 jdrewniak@deploy1003: jdrewniak, milazg: Continuing with deployment
* 20:34 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1346.eqiad.wmnet
* 20:34 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1346.eqiad.wmnet
* 20:32 jdrewniak@deploy1003: jdrewniak, milazg: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2005.codfw.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:28 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1328215{{!}}Remove mode from RestModuleOverrides (T434267)]]
* 20:27 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2005.codfw.wmnet
* 20:26 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:24 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:22 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] (duration: 11m 24s)
* 20:17 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 20:15 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2005.codfw.wmnet
* 20:14 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 20:11 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:07 jdrewniak@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.18,1.47.0-wmf.19,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted
* 20:05 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338975{{!}}Bumping portals to master (T128546)]]
* 20:02 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2005.codfw.wmnet
* 20:01 eevans@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on sessionstore2005.codfw.wmnet with reason: Firmware upgrades — [[phab:T437516|T437516]]
* 19:53 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:50 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1346.eqiad.wmnet with reason: host reimage
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1346
* 19:38 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1346
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1346.eqiad.wmnet 196.48.64.10.in-addr.arpa 6.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:38 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1346 - rzl@cumin2003"
* 19:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1346
* 19:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1346.eqiad.wmnet with OS trixie
* 19:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1346.eqiad.wmnet
* 19:32 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1346.eqiad.wmnet
* 19:21 bking@cumin2003: END (FAIL) - Cookbook sre.hadoop.roll-restart-workers (exit_code=99) restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 thcipriani: Gerrit downtime incoming for upgrade
* 19:17 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop analytics cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hadoop.roll-restart-workers (exit_code=0) restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 19:17 dzahn@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:30:00 on gerrit.wikimedia.org with reason: maintenance upgrade
* 19:16 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on gerrit2003.wikimedia.org with reason: maintenance upgrade
* 19:04 bking@cumin2003: START - Cookbook sre.hadoop.roll-restart-workers restart workers for Hadoop test cluster: Roll restart of jvm daemons for openjdk upgrade.
* 18:21 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e] (duration: 00m 59s)
* 18:20 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (thin): Regular analytics weekly train THIN [analytics/refinery@92b5c04e]
* 18:19 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e] (duration: 05m 13s)
* 18:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 18:14 otto@deploy1003: Started deploy [analytics/refinery@92b5c04]: Regular analytics weekly train [analytics/refinery@92b5c04e]
* 18:13 otto@deploy1003: Finished deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e] (duration: 00m 39s)
* 18:13 otto@deploy1003: Started deploy [analytics/refinery@92b5c04] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@92b5c04e]
* 18:13 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 18:12 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 18:11 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 18:11 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit2002.wikimedia.org with reason: maintenance upgrade
* 18:11 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:11 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 18:10 dzahn@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on gerrit1003.wikimedia.org with reason: maintenance upgrade
* 18:09 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 18:08 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 18:06 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 16:40 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 16:35 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 16:33 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]] synced to the te
* 16:28 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338984{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338990{{!}}Provider: Cache the valid configuration in the process (T437588)]], [[gerrit:1338983{{!}}refactor: Provider: Add IConfigurationProvider::invalidateCache() (T437588)]], [[gerrit:1338986{{!}}Provider: Cache the valid configuration in the process (T437588)]]
* 15:33 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 15:04 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] (duration: 08m 08s)
* 15:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host clouddumps1001.wikimedia.org with OS bookworm
* 14:59 samtar@deploy1003: samtar: Continuing with deployment
* 14:58 samtar@deploy1003: samtar: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:56 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1335032{{!}}IS: enable wgEnableWatchstarPopover on test.wikipedia (T436955)]]
* 14:40 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:21 cdanis@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: etcd fanout fix & details UX - cdanis@cumin1004
* 14:20 cdanis@cumin1004: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "etcd fanout fix & details UX - cdanis@cumin1004"
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on clouddumps1001.wikimedia.org with reason: host reimage
* 14:07 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto cluster: Roll restart of all Presto's jvm daemons.
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host clouddumps1001.wikimedia.org with OS bookworm
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:56 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:55 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:51 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:42 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:41 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:48 klausman@dns1004: END - running authdns-update
* 12:46 klausman@dns1004: START - running authdns-update
* 12:35 dbrant@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 12:35 dbrant@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 12:34 dbrant@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 12:33 dbrant@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 12:05 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning x4
* 12:01 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:55 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:54 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1144.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:52 cgoubert@dns1004: END - running authdns-update
* 11:49 cgoubert@dns1004: START - running authdns-update
* 11:31 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:25 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:24 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1204.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:16 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1200.eqiad.wmnet
* 11:16 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1200.eqiad.wmnet
* 11:04 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1200.eqiad.wmnet with reason: Upgrading RAID firmware
* 11:04 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for an-worker1199.eqiad.wmnet
* 11:03 btullis@cumin1004: START - Cookbook sre.hosts.remove-downtime for an-worker1199.eqiad.wmnet
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:50 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 10:42 btullis@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on an-worker1199.eqiad.wmnet with reason: Upgrading RAID firmware
* 10:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on clouddb1024.eqiad.wmnet with reason: Cloning x4
* 10:00 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 09:56 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 09:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1024.eqiad.wmnet
* 09:45 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 09:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 09:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:24 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 09:04 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:51 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:46 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 08:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning x4
* 08:43 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 08:41 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:34 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] (duration: 09m 56s)
* 08:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:30 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:29 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 08:27 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 08:24 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1338705{{!}}UIC: Fix getOpenContext when UserCardButton.vue is clicked (T435299)]]
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 08:23 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 08:11 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 08:09 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 07:58 XioNoX: netflow1004:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:56 XioNoX: netflow2005:~$ sudo dpkg -i gnmic_0.48.0_Linux_x86_64.deb
* 07:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 07:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:03 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:59 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 06:43 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:42 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:42 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075 cloud-private - filippo@cumin1003"
* 06:39 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 06:38 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 06:37 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 05:04 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 05:03 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tool-server' for release 'main' .
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1345.eqiad.wmnet
* 00:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1345.eqiad.wmnet
* 00:11 dzahn@dns1004: END - running authdns-update
* 00:08 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 00:08 dzahn@dns1004: START - running authdns-update
== 2026-09-09 ==
* 23:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:46 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1345.eqiad.wmnet with reason: host reimage
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:37 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1345
* 23:33 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1345
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1345.eqiad.wmnet 195.48.64.10.in-addr.arpa 5.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:33 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:32 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1345 - rzl@cumin2003"
* 23:29 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1345
* 23:28 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1345.eqiad.wmnet with OS trixie
* 23:28 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1345.eqiad.wmnet
* 23:27 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1345.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1344.eqiad.wmnet
* 23:26 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1344.eqiad.wmnet
* 23:22 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] (duration: 11m 15s)
* 23:18 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 23:16 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:15 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 23:11 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338142{{!}}Move testcommonswiki links tables to x4 (T398709)]]
* 22:57 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:51 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1344.eqiad.wmnet with reason: host reimage
* 22:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore2004.codfw.wmnet with OS bookworm
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1344
* 22:39 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1344
* 22:37 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1344
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1344.eqiad.wmnet 194.48.64.10.in-addr.arpa 4.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:37 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1344 - rzl@cumin2003"
* 22:33 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1344
* 22:33 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1344.eqiad.wmnet with OS trixie
* 22:32 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] (duration: 10m 21s)
* 22:32 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1344.eqiad.wmnet
* 22:31 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1344.eqiad.wmnet
* 22:27 derenrich@deploy1003: derenrich, egardner: Continuing with deployment
* 22:26 derenrich@deploy1003: derenrich, egardner: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:24 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:22 rzl@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1343.eqiad.wmnet
* 22:22 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1338303{{!}}Revert "Enable discord preview extension code on testwiki"]]
* 22:20 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore2004.codfw.wmnet with reason: host reimage
* 22:19 derenrich@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] (duration: 13m 40s)
* 22:16 derenrich@deploy1003: derenrich: Rolling back deployment
* 22:10 derenrich@deploy1003: derenrich: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:05 derenrich@deploy1003: Started scap sync-world: Backport for [[gerrit:1337999{{!}}Enable discord preview extension code on testwiki (T437344)]]
* 22:03 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:44 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:40 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1343.eqiad.wmnet with reason: host reimage
* 21:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:36 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1343
* 21:28 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1343
* 21:27 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1343
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1343.eqiad.wmnet 193.48.64.10.in-addr.arpa 3.9.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:27 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:27 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1343 - rzl@cumin2003"
* 21:23 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1343
* 21:22 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] (duration: 12m 29s)
* 21:21 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1343.eqiad.wmnet with OS trixie
* 21:21 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1343.eqiad.wmnet
* 21:20 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1343.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1342.eqiad.wmnet
* 21:19 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1342.eqiad.wmnet
* 21:17 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1075.eqiad.wmnet with OS trixie
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:14 jforrester@deploy1003: jforrester: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:09 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 21:09 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1338273{{!}}wikifunctions: Move client fragments to mainstash (T432849)]]
* 21:08 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host sessionstore2004.codfw.wmnet with OS bookworm
* 21:07 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 21:07 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:06 eevans@cumin1003: START - Cookbook sre.hosts.provision for host sessionstore2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 21:05 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sessionstore2004.codfw.wmnet
* 20:59 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:57 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-canary cluster: Roll restart of all Presto's jvm daemons.
* 20:56 bking@cumin2003: END (PASS) - Cookbook sre.presto.roll-restart-workers (exit_code=0) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:54 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host sessionstore2004.codfw.wmnet
* 20:53 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:53 bking@cumin2003: END (ERROR) - Cookbook sre.presto.roll-restart-workers (exit_code=97) for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:53 bking@cumin2003: START - Cookbook sre.presto.roll-restart-workers for Presto an-presto-test cluster: Roll restart of all Presto's jvm daemons.
* 20:50 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1075.eqiad.wmnet with reason: host reimage
* 20:49 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* {{safesubst:SAL entry|1=20:45 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2}}
* 20:42 rzl@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1342.eqiad.wmnet with reason: host reimage
* 20:41 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts sessionstore2004.codfw.wmnet
* 20:40 sbassett@deploy1003: aranyap, sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:39 sbassett@deploy1003: aranyap, sbassett: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "Filter}}
* 20:35 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* {{safesubst:SAL entry|1=20:34 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338280{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338281{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338283{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338282{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338278{{!}}Revert^2 "Filter out non-http(s) license urls"]], [[gerrit:1338279{{!}}Revert^2 "}}
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1342
* 20:29 rzl@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1342
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1342.eqiad.wmnet 160.32.64.10.in-addr.arpa 0.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:29 rzl@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:28 rzl@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1342 - rzl@cumin2003"
* 20:24 rzl@cumin2003: START - Cookbook sre.dns.netbox
* 20:24 rzl@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1342
* 20:23 rzl@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1342.eqiad.wmnet with OS trixie
* 20:23 rzl@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1342.eqiad.wmnet
* 20:23 rzl@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1342.eqiad.wmnet
* 20:22 rzl@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1342.eqiad.wmnet
* 20:15 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1027.eqiad.wmnet with OS bookworm
* 19:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1027.eqiad.wmnet with reason: host reimage
* 19:41 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1027.eqiad.wmnet with OS bookworm
* 19:36 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:35 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:28 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1027.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:26 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:19 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1075.eqiad.wmnet with OS trixie
* 19:19 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:06 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1075.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:05 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1075
* 19:05 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1075
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:04 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 19:03 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1075] - vriley@cumin1003"
* 18:59 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:23 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 18:06 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1026.eqiad.wmnet with OS bookworm
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 18:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 18:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:01 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 18:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 17:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 17:57 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 17:56 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:54 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 17:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 17:49 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:49 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 17:49 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 17:45 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1026.eqiad.wmnet with reason: host reimage
* 17:36 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1026.eqiad.wmnet with OS bookworm
* 17:31 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1026.eqiad.wmnet with OS bookworm
* 17:29 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:27 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 17:23 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 17:12 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 17:04 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1026.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 16:46 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] (duration: 09m 28s)
* 16:41 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 16:41 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:36 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338173{{!}}testwiki: allow testing new AccountSetup experiment (T436872)]]
* 16:28 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:25 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 15:55 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 15:54 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 15:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] (duration: 09m 43s)
* 15:41 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 15:40 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:36 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338213{{!}}Disable parsoidCachePrewarm jobs everywhere (T436001 T436206)]]
* 15:17 urbanecm: Delete all running periodic jobs starting with `growthexperiments-refreshlinkrecommendations-*` (to pick up new configuration; [[phab:T392944|T392944]])
* 15:08 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 15:07 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 15:06 moritzm: installing grub2 bugfix updates from Bookworm point release
* 15:04 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp6008.drmrs.wmnet
* 15:01 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:42 hnowlan: half concurrency for parsoidCachePrewarm RecordLintJob and refreshLinks in jobqueue, eqiad & codfw
* 14:35 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-jobrunner: apply
* 14:34 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-jobrunner: apply
* 14:32 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tool-server' for release 'main' .
* 14:20 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:07 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 14:07 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 14:06 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - sukhe@cumin1003"
* 13:55 btullis@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 13:53 btullis@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 13:43 btullis@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 13:42 btullis@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 13:29 moritzm: pruned obsolete Bullseye image dispatch from the docker registry [[phab:T416452|T416452]]
* 13:28 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:26 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b7-eqiad
* 13:25 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 13:24 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 13:22 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 13:22 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:17 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a4-eqiad
* 13:17 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] (duration: 10m 15s)
* 13:14 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 13:11 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:10 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1334945{{!}}ArticleGuidance: Add the redirect configuration keys (T434487)]], [[gerrit:1337983{{!}}Replace experiment with instrument and config-driven redirect (T434487)]]
* 13:02 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on ldap-rw[1001,2001].wikimedia.org with reason: work in progress
* 12:49 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 12:48 btullis@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 12:46 btullis@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] (duration: 14m 39s)
* 12:41 ladsgroup@deploy1003: tryvix1509, ladsgroup: Continuing with deployment
* 12:35 ladsgroup@deploy1003: tryvix1509, ladsgroup: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1338151{{!}}core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki (T437409)]]
* 12:16 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:16 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 11:53 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] (duration: 21m 58s)
* 11:48 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:35 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:31 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1338031{{!}}[Growth] Enable iterative Add Link task pool population on all wikis (T392944)]]
* 10:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Repooling db2196
* 10:47 moritzm: pruned obsolete Bullseye images nodejs12-slim/nodejs12-devel/nodejs14-slim/nodejs16-slim from the docker registry [[phab:T416452|T416452]]
* 10:43 moritzm: installing Bird security updates
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1260: Repooling after cloning
* 10:09 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Repooling db2196
* 10:07 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:06 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 09:55 moritzm: pruned obsolete Bullseye images openjdk-8-jdk/openjdk-8-jre/openjdk-11-jre/openjdk-11-jdk from the docker registry [[phab:T416452|T416452]]
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1260: Repooling after cloning
* 09:52 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 09:52 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 09:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 09:28 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:27 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 09:03 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 09:02 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 09:01 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 16 hosts with reason: upgrade ssw1-a1-eqiad
* 08:58 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 22 hosts with reason: upgrade ssw1-a1-eqiad
* 08:56 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:55 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:49 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 08:49 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 08:48 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 08:48 cmooney@cumin1004: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 08:36 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1155.eqiad.wmnet with reason: Cloning sanitarium
* 08:30 brouberol@dns1004: END - running authdns-update
* 08:28 moritzm: pruned obsolete Bullseye image golang1.15 from the docker registry [[phab:T416452|T416452]]
* 08:28 brouberol@dns1004: START - running authdns-update
* 08:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1260: Needs to clone another host from this one
* 08:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1260: Needs to clone another host from this one
* 08:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1260.eqiad.wmnet with reason: Cloning sanitarium
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:01 btullis@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 08:00 btullis@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts an-worker1198.eqiad.wmnet
* 07:40 chlod: UTC morning backport window done
* 07:37 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] (duration: 21m 36s)
* 07:32 chlod@deploy1003: chlod, hamishz: Continuing with deployment
* 07:20 chlod@deploy1003: chlod, hamishz: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1335706{{!}}thwikibooks: update tagline and wordmark (T436426)]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 45s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1025.eqiad.wmnet with OS bookworm
== 2026-09-08 ==
* 23:51 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:48 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1025.eqiad.wmnet with reason: host reimage
* 23:39 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1313.eqiad.wmnet
* 23:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1313.eqiad.wmnet
* 23:35 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:30 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:25 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 23:19 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:19 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:15 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1025.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 23:07 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 23:05 Amir1: dropped 57 tables on db1260 ([[phab:T437278|T437278]])
* 23:03 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 23:03 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 23:02 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 22:57 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1313.eqiad.wmnet with reason: host reimage
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1313
* 22:38 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1313
* 22:37 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1313
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1313.eqiad.wmnet 149.32.64.10.in-addr.arpa 9.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:37 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:37 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1313 - swfrench@cumin1003"
* 22:33 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:33 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1313
* 22:32 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1313.eqiad.wmnet with OS trixie
* 22:32 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1313.eqiad.wmnet
* 22:31 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1313.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1306.eqiad.wmnet
* 22:30 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1306.eqiad.wmnet
* 22:27 brett@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp6008.drmrs.wmnet with OS trixie
* 22:18 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 22:03 brett@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 22:01 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] (duration: 09m 53s)
* 22:00 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:59 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1025.eqiad.wmnet with OS bookworm
* 21:58 Amir1: drop links tables from db2210 ([[phab:T437278|T437278]])
* 21:57 jdrewniak@deploy1003: jdrewniak: Continuing with deployment
* 21:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:56 jdrewniak@deploy1003: jdrewniak: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:52 brett@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp6008.drmrs.wmnet with reason: host reimage
* 21:51 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338053{{!}}Bumping portals to master (T128546)]]
* 21:51 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1306.eqiad.wmnet with reason: host reimage
* 21:48 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1025.eqiad.wmnet with OS bookworm
* 21:45 jdrewniak@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] (duration: 05m 27s)
* 21:43 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Continuing with deployment
* 21:40 jdrewniak@deploy1003: portalsbuilder, jdrewniak: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 jdrewniak@deploy1003: Started scap sync-world: Backport for [[gerrit:1338049{{!}}Assets build - 2026-09-08 21:25:19+00:00]]
* 21:35 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1024.eqiad.wmnet with OS bookworm
* 21:33 brett@cumin2003: START - Cookbook sre.hosts.reimage for host cp6008.drmrs.wmnet with OS trixie
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1306
* 21:31 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1306
* 21:30 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1306
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1306.eqiad.wmnet 146.32.64.10.in-addr.arpa 6.4.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:30 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1306 - swfrench@cumin1003"
* 21:24 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] (duration: 09m 12s)
* 21:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 21:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1306
* 21:19 reedy@deploy1003: reedy: Continuing with deployment
* 21:19 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1306.eqiad.wmnet with OS trixie
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:19 reedy@deploy1003: reedy: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1306.eqiad.wmnet
* 21:18 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1306.eqiad.wmnet
* 21:15 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1024.eqiad.wmnet with reason: host reimage
* 21:15 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1324753{{!}}InitialiseSettings: Enable 2FA enforcement on remaining private wikis (T428103)]]
* {{safesubst:SAL entry|1=21:10 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out}}
* 21:05 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1024.eqiad.wmnet with OS bookworm
* 21:05 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 21:04 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=21:03 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out non-http(s) lice}}
* 20:59 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1023.eqiad.wmnet
* {{safesubst:SAL entry|1=20:58 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338037{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338034{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338035{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338036{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338032{{!}}Revert "Filter out non-http(s) license urls"]], [[gerrit:1338033{{!}}Revert "Filter out n}}
* 20:53 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 20:50 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1305.eqiad.wmnet
* 20:50 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1305.eqiad.wmnet
* 20:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 20:34 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 20:28 sbassett@deploy1003: sbassett: Continuing with deployment
* {{safesubst:SAL entry|1=20:27 sbassett@deploy1003: sbassett: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-http(s) license}}
* 20:27 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1023.eqiad.wmnet with OS bookworm
* {{safesubst:SAL entry|1=20:23 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1338023{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337997{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338024{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337998{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1338025{{!}}Filter out non-http(s) license urls (T435999)]], [[gerrit:1337996{{!}}Filter out non-}}
* 20:15 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] (duration: 10m 16s)
* 20:13 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:10 aaron@deploy1003: aaron: Continuing with deployment
* 20:09 aaron@deploy1003: aaron: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 20:05 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1305.eqiad.wmnet with reason: host reimage
* 20:05 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1326425{{!}}Add wmf-analytics-commons external module to commonswiki (T434927)]]
* 20:01 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1023.eqiad.wmnet with reason: host reimage
* 19:51 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1023.eqiad.wmnet with OS bookworm
* 19:45 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1305
* 19:44 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1305
* 19:43 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1305
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1305.eqiad.wmnet 138.32.64.10.in-addr.arpa 8.3.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:43 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:43 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1305 - swfrench@cumin1003"
* 19:39 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 19:39 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1305
* 19:38 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1305.eqiad.wmnet with OS trixie
* 19:37 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1023.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1305.eqiad.wmnet
* 19:37 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1305.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1275.eqiad.wmnet
* 19:31 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1275.eqiad.wmnet
* 19:23 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 19:17 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:12 sukhe@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:04 sukhe@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 18:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:53 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1275.eqiad.wmnet with reason: host reimage
* 18:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1275
* 18:34 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1275
* 18:33 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1275
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1275.eqiad.wmnet 168.48.64.10.in-addr.arpa 8.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:33 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:25 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1275 - swfrench@cumin1003"
* 18:20 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1275
* 18:20 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1275.eqiad.wmnet with OS trixie
* 18:19 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1275.eqiad.wmnet
* 18:19 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1275.eqiad.wmnet
* 18:18 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 17:43 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy1003.eqiad.wmnet
* 17:36 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:36 cdobbins@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-ntp (exit_code=0) rolling restart_daemons on A:dnsbox
* 17:30 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy1003.eqiad.wmnet
* 17:25 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:15 swfrench@deploy1003: Finished scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup (duration: 04m 18s)
* 17:11 Amir1: dropping links tables from db1247 (s4 replica) - ([[phab:T437278|T437278]])
* 17:10 swfrench@deploy1003: Started scap sync-world: Noop deployment to test mediawiki_runtime_image cleanup
* 16:51 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] (duration: 10m 19s)
* 16:46 zabe@deploy1003: zabe: Continuing with deployment
* 16:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337946{{!}}Use local database for category table in SpecialWantedCategories (T437286)]], [[gerrit:1337947{{!}}Use local database for category table in SpecialWantedCategories (T437286)]]
* 16:29 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 jhancock@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:25 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2005.codfw.wmnet
* 16:24 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker2004.codfw.wmnet
* 16:12 jhancock@cumin2003: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 16:08 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:05 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 16:04 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2003.codfw.wmnet
* 15:58 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:55 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2003.codfw.wmnet
* 15:54 jhancock@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:44 jhancock@cumin2003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 15:29 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2002.codfw.wmnet
* 15:19 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2002.codfw.wmnet
* 14:55 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:46 logmsgbot: dreamyjazz Deployed security patch for [[phab:T436429|T436429]]
* 14:45 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cephosd2001.codfw.wmnet
* 14:44 topranks: shutdown et-1/1/5 on cr1-codfw to shift traffic off ssw1-a1-codfw
* 14:43 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: upgrade ssw1-a1-eqiad
* 14:34 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host cephosd2001.codfw.wmnet
* 14:33 btullis@cumin1003: END (ERROR) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=97) rolling reboot on A:cephosd-codfw
* 14:30 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 14:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 14:28 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 14:22 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 14:13 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --olderThan {{Gerrit|1788220800}} --verbose # [[phab:T437158|T437158]]
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 14:12 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 14:12 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 14:03 topranks: drain traffic from ssw1-a1-codfw before JunOS upgrade [[phab:T426197|T426197]]
* 14:02 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 13:58 cgoubert@deploy1003: helmfile [staging-codfw] DONE helmfile.d/services/mw-debug: apply
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 13:57 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 13:57 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 13:56 cgoubert@deploy1003: helmfile [staging-codfw] START helmfile.d/services/mw-debug: apply
* 13:55 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:54 stran@deploy1003: mwscript-k8s job started: foreachwikiindblist checkuser-suggested-investigations extensions/CheckUser/maintenance/populateSiCaseProperties.php # [[phab:T435066|T435066]]
* 13:52 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:51 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 13:50 cgoubert@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/services/mw-debug: apply
* 13:46 cdobbins@cumin1003: START - Cookbook sre.dns.roll-restart-ntp rolling restart_daemons on A:dnsbox
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 13:46 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 13:46 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 13:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2012.codfw.wmnet with OS bookworm
* 13:44 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] (duration: 34m 00s)
* 13:40 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 13:37 cgoubert@deploy1003: helmfile [staging-eqiad] START helmfile.d/services/mw-debug: apply
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 13:35 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 13:35 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 13:32 stran@deploy1003: stran: Continuing with deployment
* 13:29 stran@deploy1003: stran: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:28 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 13:23 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 13:23 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2012.codfw.wmnet with reason: host reimage
* 13:21 moritzm: installing qemu security updates
* 13:18 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 13:16 ayounsi@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'email' for AS: 139628
* 13:15 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 139628
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 13:13 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 13:13 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 13:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 2519
* 13:11 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 2519
* 13:10 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 14593
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:10 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337610{{!}}Split out edit and block-based filters from activity filters (T436508)]]
* 13:09 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:08 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 14593
* 13:06 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2012.codfw.wmnet with OS bookworm
* 13:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2012.codfw.wmnet
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2012.codfw.wmnet
* 13:05 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:04 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2012.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 13:01 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 13:01 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 12:58 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 34655
* 12:58 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'email' for AS: 34655
* 12:56 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2012.codfw.wmnet
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'clear' for AS: 35320
* 12:55 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'clear' for AS: 35320
* 12:55 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:55 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:55 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-codfw
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-codfw
* 12:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2011.codfw.wmnet with OS bookworm
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-codfw
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-by27-esams
* 12:54 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-by27-esams
* 12:54 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:54 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b13-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-b12-drmrs
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-esams
* 12:53 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-esams
* 12:53 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device asw1-bw27-esams
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-drmrs
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-esams
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-esams
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-drmrs
* 12:51 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-eqsin
* 12:51 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-eqsin
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr4-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr3-ulsfo
* 12:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr3-ulsfo
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f3-eqiad
* 12:50 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 12:50 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f3-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-eqiad
* 12:50 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e3-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-f4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-c8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-c8-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-e4-eqiad
* 12:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e1-eqiad
* 12:44 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e1-eqiad
* 12:43 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 12:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-b1-codfw
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr1-eqiad
* 12:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-eqiad
* 12:42 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cr2-eqiad
* 12:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f1-eqiad
* 12:41 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f1-eqiad
* 12:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cloudsw1-d5-eqiad
* 12:40 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device cloudsw1-d5-eqiad
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 12:38 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 12:38 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 12:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:32 klausman@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 12:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2011.codfw.wmnet with reason: host reimage
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 12:22 klausman@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 12:11 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2012.codfw.wmnet
* 12:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2011.codfw.wmnet with OS bookworm
* 12:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2011.codfw.wmnet
* 12:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2011.codfw.wmnet
* 12:09 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2011.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:06 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] (duration: 09m 54s)
* 12:01 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2011.codfw.wmnet
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Continuing with deployment
* 12:01 zabe@deploy1003: zabe, somerandomdeveloper: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2010.codfw.wmnet with OS bookworm
* 11:56 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337898{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]], [[gerrit:1337899{{!}}Skip DB read-only check for CentralAuthSessionManager (T437273)]]
* 11:46 marostegui@dns1004: END - running authdns-update
* 11:44 marostegui@dns1004: START - running authdns-update
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:40 Amir1: dropping unneeded tables from x4 - db1260 ([[phab:T437278|T437278]])
* 11:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2010.codfw.wmnet with reason: host reimage
* 11:23 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2011.codfw.wmnet
* 11:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2010.codfw.wmnet with OS bookworm
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2010.codfw.wmnet
* 11:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:21 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2010.codfw.wmnet
* 11:20 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2010.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2010.codfw.wmnet
* 11:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2009.codfw.wmnet with OS bookworm
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:57 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:56 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2009.codfw.wmnet with reason: host reimage
* 10:43 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] (duration: 10m 57s)
* 10:39 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:38 samtar@deploy1003: samtar: Continuing with deployment
* 10:37 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:37 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:36 samtar@deploy1003: samtar: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:34 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:32 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307862{{!}}IS/IS-labs: wgEnableWatchstarPopover default false, enable for enwiki beta (T431355)]]
* 10:30 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2010.codfw.wmnet
* 10:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2009.codfw.wmnet with OS bookworm
* 10:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2009.codfw.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2009.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:18 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2009.codfw.wmnet
* 10:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2008.codfw.wmnet with OS bookworm
* 10:07 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:05 mvernon@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 10:04 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 10:01 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:59 mvernon@cumin1003: END (ERROR) - Cookbook sre.hardware.upgrade-firmware (exit_code=97) upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2008.codfw.wmnet with reason: host reimage
* 09:45 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] (duration: 13m 15s)
* 09:45 ayounsi@dns1004: END - running authdns-update
* 09:44 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2009.codfw.wmnet
* 09:43 ayounsi@dns1004: START - running authdns-update
* 09:39 zabe@deploy1003: zabe: Continuing with deployment
* 09:37 zabe@deploy1003: zabe: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: remove include for former GRE tunnels v6 PTR - ayounsi@cumin1003"
* 09:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2008.codfw.wmnet with OS bookworm
* 09:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:32 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337856{{!}}Use virtual domain in NameTableStore for collation (T405812)]], [[gerrit:1337857{{!}}Use virtual domain in NameTableStore for collation (T405812)]]
* 09:32 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2008.codfw.wmnet
* 09:31 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2008.codfw.wmnet
* 09:31 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2008.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 09:29 XioNoX: remove GRE tunnels eqiad-drmrs eqdfw-ulsfo
* 09:23 moritzm: installing rsync security updates
* 09:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2007.codfw.wmnet with OS bookworm
* 09:22 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2008.codfw.wmnet
* 09:11 marostegui@cumin1003: dbctl commit (dc=all): 'Make x4 and s4 RW again [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96393 and previous config saved to /var/cache/conftool/dbconfig/20260908-091121-marostegui.json
* 09:07 marostegui@cumin1003: dbctl commit (dc=all): 'Remove old s4 masters from x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96392 and previous config saved to /var/cache/conftool/dbconfig/20260908-090749-marostegui.json
* 09:05 marostegui@cumin1003: dbctl commit (dc=all): 'Set x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96391 and previous config saved to /var/cache/conftool/dbconfig/20260908-090517-marostegui.json
* 09:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Set s4 commons to read-only for maintenance [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96389 and previous config saved to /var/cache/conftool/dbconfig/20260908-090228-marostegui.json
* 09:02 marostegui: Starting x4 split from s4, RO time on commons needed [[phab:T404715|T404715]]
* 09:00 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2007.codfw.wmnet with reason: host reimage
* 08:58 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2008.codfw.wmnet
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:43 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 32 hosts with reason: x4 split
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2007.codfw.wmnet with OS bookworm
* 08:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:37 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2007.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:37 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin2003.codfw.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2007.codfw.wmnet
* 08:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2007.codfw.wmnet
* 08:35 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:34 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1003.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:29 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: Release v0.11.3 - ayounsi@cumin1003
* 08:26 mvernon@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs2007.codfw.wmnet
* 08:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2006.codfw.wmnet with OS bookworm
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 07:58 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2006.codfw.wmnet with reason: host reimage
* 07:50 mvernon@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2007.codfw.wmnet
* 07:43 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2006.codfw.wmnet with OS bookworm
* 07:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2006.codfw.wmnet
* 07:37 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 07:30 denisse: Add grafana-plugins 0.15 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 07:29 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:27 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 07:27 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:27 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:26 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:26 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 07:22 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 07:22 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2006.codfw.wmnet
* 07:14 jmm@dns1004: END - running authdns-update
* 07:13 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96388 and previous config saved to /var/cache/conftool/dbconfig/20260908-071308-marostegui.json
* 07:12 jmm@dns1004: START - running authdns-update
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96387 and previous config saved to /var/cache/conftool/dbconfig/20260908-071216-marostegui.json
* 07:12 marostegui@cumin1003: dbctl commit (dc=all): 'Remove hosts from s4 as they should only be in x4 [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P96386 and previous config saved to /var/cache/conftool/dbconfig/20260908-071159-marostegui.json
* 05:07 denisse: Add grafana-plugins 0.10 to bookworm-wikimedia - [[phab:T436056|T436056]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.16 (duration: 02m 27s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]] (duration: 36m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.19 refs [[phab:T430838|T430838]]
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 41s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-07 ==
* 21:52 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] (duration: 11m 00s)
* 21:47 zabe@deploy1003: zabe: Continuing with deployment
* 21:45 zabe@deploy1003: zabe: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337635{{!}}Move globalimagelinks queries to x4 (T437108)]]
* 21:37 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] (duration: 09m 34s)
* 21:33 zabe@deploy1003: zabe: Continuing with deployment
* 21:32 zabe@deploy1003: zabe: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337319{{!}}Move production reads for commons link tables to x4 for all requests (T437108)]]
* 21:03 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] (duration: 10m 27s)
* 20:58 zabe@deploy1003: zabe: Continuing with deployment
* 20:57 zabe@deploy1003: zabe: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:52 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337656{{!}}Move production reads for commons link tables to x4 for 50% of requests (T437108)]]
* 20:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set weight of db1261 to zero in s4 ([[phab:T437108|T437108]])', diff saved to https://phabricator.wikimedia.org/P96385 and previous config saved to /var/cache/conftool/dbconfig/20260907-203804-ladsgroup.json
* 20:23 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] (duration: 09m 28s)
* 20:19 zabe@deploy1003: zabe: Continuing with deployment
* 20:18 zabe@deploy1003: zabe: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337655{{!}}Move production reads for commons link tables to x4 for 25% of requests (T437108)]]
* 20:12 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] (duration: 10m 06s)
* 20:07 zabe@deploy1003: zabe, daimona: Continuing with deployment
* 20:06 zabe@deploy1003: zabe, daimona: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:02 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1326400{{!}}Stop setting wgCampaignEventsEnableWorklists (T429510)]]
* 19:59 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] (duration: 11m 27s)
* 19:55 zabe@deploy1003: zabe: Continuing with deployment
* 19:52 zabe@deploy1003: zabe: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:48 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337650{{!}}Move production reads for commons link tables to x4 for 10% of requests (T437108)]]
* 19:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] (duration: 11m 40s)
* 19:26 zabe@deploy1003: zabe: Continuing with deployment
* 19:23 zabe@deploy1003: zabe: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:19 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336120{{!}}Move production reads for commons link tables to x4 for 1% of requests (T437108)]]
* 19:07 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] (duration: 14m 17s)
* 19:00 zabe@deploy1003: zabe: Continuing with deployment
* 18:57 zabe@deploy1003: zabe: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:53 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1328365{{!}}Set RemoteVirtualDomainsMapping for test-commons]]
* 18:33 zabe@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:32 zabe@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:31 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] (duration: 09m 12s)
* 18:27 zabe@deploy1003: zabe: Continuing with deployment
* 18:26 zabe@deploy1003: zabe: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:22 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337646{{!}}SpecialMostGloballyLinkedFiles: Recache from the globalusage domain (T400824)]]
* 16:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2005.codfw.wmnet with OS bookworm
* 16:01 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] (duration: 10m 22s)
* 15:59 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:57 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:56 zabe@deploy1003: zabe: Continuing with deployment
* 15:55 zabe@deploy1003: zabe: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:51 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1336829{{!}}Disable query pages not yet compatible with commons split (T437108)]]
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:47 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:46 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2005.codfw.wmnet with reason: host reimage
* 15:44 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:44 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:27 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2005.codfw.wmnet with OS bookworm
* 15:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2005.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2005.codfw.wmnet
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:14 mvernon@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:11 moritzm: installing rsync security updates
* 15:04 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1341.eqiad.wmnet
* 15:04 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1341.eqiad.wmnet
* 15:03 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2005.codfw.wmnet
* 15:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2004.codfw.wmnet with OS bookworm
* 15:00 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:58 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 14:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 14:41 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2004.codfw.wmnet with reason: host reimage
* 14:40 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:37 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 14:37 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 14:35 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:35 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 14:30 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:29 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:25 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 14:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1228: Repooling db1228 into s4
* 14:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2004.codfw.wmnet with OS bookworm
* 14:21 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Repooling after cloning
* 14:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:19 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2004.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2004.codfw.wmnet
* 14:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2004.codfw.wmnet
* 14:18 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 14:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:14 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1074.eqiad.wmnet
* 14:14 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 14:13 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1341.eqiad.wmnet with reason: host reimage
* 14:13 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* {{safesubst:SAL entry|1=14:11 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mul}}
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2004.codfw.wmnet
* 14:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1074.eqiad.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1073.eqiad.wmnet
* 14:07 krinkle@deploy1003: krinkle: Continuing with deployment
* {{safesubst:SAL entry|1=14:04 krinkle@deploy1003: krinkle: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with multiple properties}}
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1073.eqiad.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1072.eqiad.wmnet
* 14:01 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1341
* 14:01 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1341
* 14:01 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:00 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1341
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1341.eqiad.wmnet 159.32.64.10.in-addr.arpa 9.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:00 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* 13:59 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1341 - cgoubert@cumin2003"
* {{safesubst:SAL entry|1=13:59 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1335754{{!}}tests: Add user_id and actor in testLogDatabaseRowsForHiddenUser (T436971)]], [[gerrit:1336613{{!}}User: Use ActorStore on User::load (T428517)]], [[gerrit:1335752{{!}}Fix empty bases in m(under{{!}}over) (T436876)]], [[gerrit:1336611{{!}}Use Core-compatible output for cancellation (T379359)]], [[gerrit:1336609{{!}}Page: Fix absence caching when combined with mult}}
* 13:59 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2004.codfw.wmnet
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2003.codfw.wmnet with OS bookworm
* 13:55 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 13:55 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1072.eqiad.wmnet
* 13:55 filippo@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudvirt1067.eqiad.wmnet
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1341
* 13:52 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1341.eqiad.wmnet with OS trixie
* 13:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1341.eqiad.wmnet
* 13:50 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1341.eqiad.wmnet
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2008.wikimedia.org
* 13:44 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2008.wikimedia.org with OS trixie
* 13:39 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1067.eqiad.wmnet
* 13:39 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1066.eqiad.wmnet
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:39 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:37 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1228: Repooling db1228 into s4
* 13:36 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Repooling after cloning
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2003.codfw.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1066.eqiad.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1065.eqiad.wmnet
* 13:28 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:27 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudvirt1065.eqiad.wmnet
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 13:26 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:25 moritzm: installing openssh security updates
* 13:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:24 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1340.eqiad.wmnet
* 13:24 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1340.eqiad.wmnet
* 13:23 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2008.wikimedia.org with reason: host reimage
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] (duration: 10m 06s)
* 13:17 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2003.codfw.wmnet with OS bookworm
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:16 stran@deploy1003: stran: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:16 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 13:15 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2003.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 13:15 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 13:12 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1337450{{!}}SI: Use new InfoChip text class in case status updater (T437020)]]
* 13:11 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2003.codfw.wmnet
* 13:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2003.codfw.wmnet
* 13:03 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2008.wikimedia.org with OS trixie
* 13:03 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:03 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 13:02 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2008.wikimedia.org on all recursors
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:02 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:02 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2008.wikimedia.org - jmm@cumin1004"
* 13:00 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2003.codfw.wmnet
* 12:58 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:54 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 12:54 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2008.wikimedia.org
* 12:47 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2003.codfw.wmnet
* 12:46 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2002.codfw.wmnet with OS bookworm
* 12:29 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Pool db1181.eqiad.wmnet in after cloning
* 12:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2002.codfw.wmnet with reason: host reimage
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica2007.wikimedia.org
* 12:22 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica2007.wikimedia.org with OS trixie
* 12:14 elukey: moved most of the Docker Registry's prefixes to a new internal S3 backend. For any docker pull failure that worked in the past, please ping me or drop a note in [[phab:T435499|T435499]] or contact the oncall SREs
* 12:07 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2002.codfw.wmnet with OS bookworm
* 12:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:02 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica2007.wikimedia.org with reason: host reimage
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2002.codfw.wmnet
* 11:54 jmm@dns1004: END - running authdns-update
* 11:52 jmm@dns1004: START - running authdns-update
* 11:46 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2002.codfw.wmnet
* 11:46 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica2007.wikimedia.org with OS trixie
* 11:46 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:46 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica2007.wikimedia.org on all recursors
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:45 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:45 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica2007.wikimedia.org - jmm@cumin1004"
* 11:41 moritzm: installing bash updates from bookworm point release
* 11:39 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:39 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica2007.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts ldap-replica1006.wikimedia.org
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:39 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: ldap-replica1006.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin1004"
* 11:35 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2002.codfw.wmnet
* 11:32 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 11:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs2001.codfw.wmnet with OS bookworm
* 11:28 jmm@cumin1004: START - Cookbook sre.hosts.decommission for hosts ldap-replica1006.wikimedia.org
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 11:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 11:18 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] (duration: 14m 08s)
* 11:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 11:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 11:11 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:11 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:10 zabe@deploy1003: zabe: Continuing with deployment
* 11:10 zabe@deploy1003: zabe: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 11:07 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 11:07 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 11:07 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 11:06 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:06 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:04 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1337550{{!}}rdbms: Discourage setting db domain to false for remote virtual domains (T422940)]], [[gerrit:1337310{{!}}Migrate querying categorylinks to virtual domain (T405812)]]
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs2001.codfw.wmnet with reason: host reimage
* 10:58 btullis@deploy1003: Finished scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]] (duration: 35m 20s)
* 10:54 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:53 jmm@dns1004: END - running authdns-update
* 10:51 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1340.eqiad.wmnet with reason: host reimage
* 10:50 jmm@dns1004: START - running authdns-update
* 10:47 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=frwiki # [[phab:T436659|T436659]]
* 10:40 urbanecm@deploy1003: mwscript-k8s job started: extensions/GrowthExperiments/maintenance/cleanMentorList.php --wiki=hrwiki # [[phab:T436659|T436659]]
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1340
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1340.eqiad.wmnet 164.32.64.10.in-addr.arpa 4.6.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:39 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1190: Depool db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1181: Depool db1181.eqiad.wmnet to then clone it to db1190.eqiad.wmnet - marostegui@cumin1003
* 10:38 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1181.eqiad.wmnet onto db1190.eqiad.wmnet
* 10:37 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1340 - cgoubert@cumin2003"
* 10:34 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Pool db1241.eqiad.wmnet in after cloning
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1006.wikimedia.org
* 10:33 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1006.wikimedia.org with OS trixie
* 10:33 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs2001.codfw.wmnet with OS bookworm
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1340
* 10:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:27 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1340.eqiad.wmnet with OS trixie
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1340.eqiad.wmnet
* 10:26 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1340.eqiad.wmnet
* 10:26 btullis@deploy1003: Started scap sync-world: Attempting to update mediawiki-cli for [[phab:T436913|T436913]]
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host aqs2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs2001.codfw.wmnet
* 10:18 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reboot-single for host aqs2001.codfw.wmnet
* 10:13 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 10:12 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1006.wikimedia.org with reason: host reimage
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:11 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:10 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:09 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:08 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:07 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:07 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:06 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:02 mvernon@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs2001.codfw.wmnet
* 10:00 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:59 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:57 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1006.wikimedia.org with OS trixie
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1006.wikimedia.org on all recursors
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/services/miscweb: apply
* 09:56 aokoth@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/services/miscweb: apply
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:54 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:53 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1006.wikimedia.org - jmm@cumin1004"
* 09:52 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:52 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] (duration: 10m 11s)
* 09:51 aokoth@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 09:50 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:49 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 09:49 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1006.wikimedia.org
* 09:48 aokoth@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:48 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 09:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 09:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:45 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 09:44 aokoth@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 09:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1337379{{!}}Enable thumb.wikimedia.org everywhere (T427465)]]
* 09:41 aokoth@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 09:38 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:37 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:34 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 09:25 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Repooling after switchover
* 09:23 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 09:15 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] (duration: 20m 12s)
* 09:12 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 139009
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host ldap-replica1005.wikimedia.org
* 09:10 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-replica1005.wikimedia.org with OS trixie
* 09:10 moritzm: rebuild software RAID following disk replacement [[phab:T437036|T437036]]
* 09:10 ayounsi@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 139009
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1022.eqiad.wmnet with OS bookworm
* 09:08 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 09:04 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 09:03 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 09:02 moritzm: installing giflib security updates
* 09:01 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:59 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:57 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:56 jmm@cumin1004: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 08:56 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:55 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-liftwing-studio: sync
* 08:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1337455{{!}}Mentorship: Don't renew mentors' away status on every cleaner run (T436659)]], [[gerrit:1337458{{!}}Ignore PersonalDashboard stubs (T428679)]]
* 08:52 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:52 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-replica1005.wikimedia.org with reason: host reimage
* 08:49 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Pool db1180.eqiad.wmnet in after cloning
* 08:48 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1022.eqiad.wmnet with reason: host reimage
* 08:42 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:41 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:40 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:40 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 08:39 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96355 and previous config saved to /var/cache/conftool/dbconfig/20260907-083904-marostegui.json
* 08:38 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host ldap-replica1005.wikimedia.org with OS trixie
* 08:38 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2157 to s5 primary [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96354 and previous config saved to /var/cache/conftool/dbconfig/20260907-083825-marostegui.json
* 08:38 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:38 marostegui: Starting s5 codfw failover from db2213 to db2157 - [[phab:T437188|T437188]]
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache ldap-replica1005.wikimedia.org on all recursors
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:37 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:37 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM ldap-replica1005.wikimedia.org - jmm@cumin1004"
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2157 with weight 0 [[phab:T437188|T437188]]', diff saved to https://phabricator.wikimedia.org/P96353 and previous config saved to /var/cache/conftool/dbconfig/20260907-083448-marostegui.json
* 08:34 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s5 [[phab:T437188|T437188]]
* 08:28 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:28 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host ldap-replica1005.wikimedia.org
* 08:22 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:20 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1022.eqiad.wmnet with OS bookworm
* 08:03 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 08:02 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 08:02 dpogorzelski@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 08:00 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 07:57 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1180.eqiad.wmnet onto db1241.eqiad.wmnet
* 07:56 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 12:00:00 on db1241.eqiad.wmnet with reason: Cloning
* 07:55 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1241: Cloning
* 07:55 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1241: Cloning
* 07:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Cloning
* 07:54 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Cloning
* 07:51 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 07:50 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 07:47 kartik@deploy1003: Finished scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] (duration: 41m 51s)
* 07:46 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 07:45 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 07:34 kartik@deploy1003: abi, kartik: Continuing with deployment
* 07:23 kartik@deploy1003: abi, kartik: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 kartik@deploy1003: Started scap sync-world: Backport for [[gerrit:1329314{{!}}ArticleGuidance: Configure feedback links to local talk pages (T433483)]]
* 06:14 moritzm: installing Chromium security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 08m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-06 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 25s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-05 ==
* 02:00 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-04 ==
* 22:07 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 21:48 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:42 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 21:30 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:42 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 20:40 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 19:27 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:19 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:13 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:12 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host sretest2013.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host sretest2013
* 19:11 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host sretest2013
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jhancock@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:11 jhancock@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding sretest2013 to codfw - jhancock@cumin1003"
* 19:07 jhancock@cumin1003: START - Cookbook sre.dns.netbox
* 18:18 inflatador: bking@clouddumps100[12] `systemctl reset-failed` to quash alerts until https://w.wiki/UBje . The systemd timer should try again tomorrow
* 17:27 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b8-eqiad
* 17:27 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b8-eqiad
* 16:37 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:36 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:35 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:33 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-b7-eqiad
* 16:33 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b7-eqiad
* 16:05 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b6-eqiad
* 16:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b6-eqiad
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:52 cgoubert@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 15:50 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:49 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:47 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing dns entries lsw1-b5-eqiad - cmooney@cumin1004"
* 15:43 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 15:10 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b5-eqiad
* 15:09 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b5-eqiad
* 14:46 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1045.eqiad.wmnet
* 14:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:40 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1003.eqiad.wmnet with OS trixie
* 14:38 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 14:37 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b4-eqiad
* 14:37 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b4-eqiad
* 14:32 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:32 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:31 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:28 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:28 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 14:26 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:24 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:21 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b3-eqiad
* 14:21 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b3-eqiad
* 14:17 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1003.eqiad.wmnet with reason: host reimage
* 14:17 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS trixie
* 14:04 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1003.eqiad.wmnet with OS trixie
* 13:54 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b2-eqiad
* 13:53 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b2-eqiad
* 13:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1002.eqiad.wmnet with OS trixie
* 13:18 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a4-eqiad
* 13:12 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1002.eqiad.wmnet with reason: host reimage
* 13:12 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a4-eqiad
* 13:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-b1-eqiad
* 13:05 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-b1-eqiad
* 12:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1002.eqiad.wmnet with OS trixie
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow3004.esams.wmnet with OS trixie
* 12:33 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:28 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow3004.esams.wmnet with reason: host reimage
* 12:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 12:15 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 12:11 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a7-eqiad
* 12:11 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a7-eqiad
* 12:01 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow3004.esams.wmnet with OS trixie
* 11:47 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a6-eqiad
* 11:47 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a6-eqiad
* 11:36 jclark@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki2003.codfw.wmnet
* 11:32 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki2003.codfw.wmnet with OS trixie
* 11:19 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:13 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki2003.codfw.wmnet with reason: host reimage
* 11:06 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a5-eqiad
* 11:06 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a5-eqiad
* 10:52 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki2003.codfw.wmnet with OS trixie
* 10:50 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:50 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki2003.codfw.wmnet on all recursors
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:49 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:49 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki2003.codfw.wmnet - jmm@cumin1004"
* 10:44 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 10:44 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki2003.codfw.wmnet
* 10:29 btullis@deploy1003: Finished scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli (duration: 41m 14s)
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:00 marostegui@cumin1003: Removing db1182 from zarcillo [[phab:T434869|T434869]]
* 10:00 marostegui@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts db1182.eqiad.wmnet
* 10:00 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:57 btullis@deploy1003: Started scap sync-world: Incorporating changes from https://gerrit.wikimedia.org/r/c/operations/dumps/+/1334846 to mediawiki-cli
* 09:53 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:53 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.decommission (exit_code=1)
* 09:53 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1182.eqiad.wmnet
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:50 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1182.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 09:46 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 09:45 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1182 from dbctl [[phab:T434869|T434869]]', diff saved to https://phabricator.wikimedia.org/P96346 and previous config saved to /var/cache/conftool/dbconfig/20260904-094527-marostegui.json
* 09:41 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1182.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host pki1003.eqiad.wmnet
* 09:39 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host pki1003.eqiad.wmnet with OS trixie
* 09:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1182: Decommissioning
* 09:31 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1182: Decommissioning
* 09:23 jmm@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 09:17 jmm@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on pki1003.eqiad.wmnet with reason: host reimage
* 09:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow5003.eqsin.wmnet with OS trixie
* 09:02 jmm@cumin1004: START - Cookbook sre.hosts.reimage for host pki1003.eqiad.wmnet with OS trixie
* 09:00 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 09:00 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: START - Cookbook sre.dns.wipe-cache pki1003.eqiad.wmnet on all recursors
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:59 jmm@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:59 jmm@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM pki1003.eqiad.wmnet - jmm@cumin1004"
* 08:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:55 jmm@cumin1004: START - Cookbook sre.dns.netbox
* 08:55 jmm@cumin1004: START - Cookbook sre.ganeti.makevm for new host pki1003.eqiad.wmnet
* 08:54 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2003.codfw.wmnet with reason: host reimage
* 08:48 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:45 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow5003.eqsin.wmnet with reason: host reimage
* 08:40 btullis@deploy1003: Finished scap sync-world: Trying again for [[phab:T436913|T436913]] (duration: 34m 26s)
* 08:35 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2003.codfw.wmnet with OS bookworm
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2003.codfw.wmnet
* 08:33 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw2001.wikimedia.org with OS trixie
* 08:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on db2196.codfw.wmnet with reason: Host crashed
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2003.codfw.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:20 elukey@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:12 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw2001.wikimedia.org with reason: host reimage
* 08:08 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2003.codfw.wmnet
* 08:07 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 08:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2002.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:03 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 08:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 07:57 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:56 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow5003.eqsin.wmnet with OS trixie
* 07:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:55 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw2001.wikimedia.org with OS trixie
* 07:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:51 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:50 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cassandra-dev2001.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:13 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ldap-rw1001.wikimedia.org with OS trixie
* 07:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2196: down
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2196: down
* 06:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:52 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ldap-rw1001.wikimedia.org with reason: host reimage
* 06:40 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host ldap-rw1001.wikimedia.org with OS trixie
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:28 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:28 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065 cloud-private - filippo@cumin1003"
* 06:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:21 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 06:20 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 00m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 01:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:02 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1074.eqiad.wmnet with OS trixie
* 01:00 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:59 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 00:46 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 00:44 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:39 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1074.eqiad.wmnet with reason: host reimage
* 00:23 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1074.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1073.eqiad.wmnet with OS trixie
* 00:23 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
== 2026-09-03 ==
* 21:46 tsev@deploy1003: mwscript-k8s job started: purgeList.php # [[phab:T435363|T435363]]
* 21:03 eevans@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:55 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] (duration: 12m 24s)
* 20:52 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:52 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1021.eqiad.wmnet with OS bookworm
* 20:50 arlolra@deploy1003: arlolra, tsev: Continuing with deployment
* 20:46 arlolra@deploy1003: arlolra, tsev: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:44 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1047.eqiad.wmnet
* 20:42 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1335109{{!}}Add exclusions to Apple app site association file (T435363)]]
* 20:41 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a3-eqiad
* 20:40 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a3-eqiad
* 20:39 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] (duration: 10m 23s)
* 20:34 arlolra@deploy1003: arlolra, jgiannelos: Continuing with deployment
* 20:33 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1047.eqiad.wmnet
* 20:33 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:32 arlolra@deploy1003: arlolra, jgiannelos: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:31 andrew@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd1046.eqiad.wmnet
* 20:28 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1021.eqiad.wmnet with reason: host reimage
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1334777{{!}}prv: Enable parsoid rendering for more wikisource wikis (T436919)]]
* 20:23 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] (duration: 13m 41s)
* 20:20 andrew@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1046.eqiad.wmnet
* 20:16 catrope@deploy1003: catrope: Continuing with deployment
* 20:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1021.eqiad.wmnet with OS bookworm
* 20:13 catrope@deploy1003: catrope: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts aqs1021.eqiad.wmnet
* 20:11 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aqs1021.eqiad.wmnet
* 20:09 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1334268{{!}}Email confirmation A/A test: make registration cutoff consistent (T435135)]], [[gerrit:1335073{{!}}Instrumentation for email confirmation upfront enforcement A/A test (T435135)]], [[gerrit:1335074{{!}}Email confirmation A/A: check creation wiki, centralize logic (T435135)]]
* 20:00 eevans@cumin1003: START - Cookbook sre.hosts.reboot-single for host aqs1021.eqiad.wmnet
* 19:49 eevans@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts aqs1021.eqiad.wmnet
* 19:19 swfrench@deploy1003: Finished scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]] (duration: 02m 59s)
* 19:16 swfrench@deploy1003: Started scap sync-world: Noop deployment to validate pretrain logstash check configuration - [[phab:T435419|T435419]]
* 19:01 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:01 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 19:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 18:57 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:57 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 18:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:34 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 18:20 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 17:55 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:55 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:54 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:53 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:52 andrew@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:46 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 17:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 17:46 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 17:45 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 17:44 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 17:44 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:40 andrew@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 17:39 ryankemper: [WDQS] Service looks healthy again, CPU load and thread count have dropped considerably over the last hour
* 17:39 ryankemper: [[phab:T421642|T421642]] [WDQS] requestctl changes: `2026-09-03 16:23-17:33` UTC: added hard-deny pair `cache-text/wdqs_futile_sparql_sep_2026_deny(+_bots)`; extended pattern `ua/wdqs_heavy_sparql_bots_2026` and added default-scope twin `wdqs_heavy_sparql_bots_jul_2026_ratelimit_default`; added ipblock `abuse/wdqs_sparql_scanners_sep_2026` + throttle `wdqs_sparql_scanners_sep_2026_ratelimit` (needed manual `requestctl update-provenance-map`)
* 17:37 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:37 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:36 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:31 andrew@cumin2003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host cloudcephosd1045.eqiad.wmnet
* 17:30 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:30 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:28 dancy@deploy1003: Installation of scap version "4.289.0" completed for 3 hosts
* 17:26 dancy@deploy1003: Installing scap version "4.289.0" for 3 host(s)
* 17:24 cmooney@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-a2-eqiad
* 17:24 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:22 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:22 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:15 cmooney@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-a2-eqiad
* 17:14 cmooney@cumin1004: START - Cookbook sre.network.tls for network device lsw1-a2-eqiad
* 17:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:10 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:10 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:09 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 17:01 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:59 andrew@cumin2003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd1045.eqiad.wmnet
* 16:58 dancy: Running scap clean-images on deploy1003
* 16:52 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:51 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:50 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 16:39 btullis@deploy1003: Started scap sync-world: Trying again for [[phab:T436913|T436913]]
* 16:14 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs2021.codfw.wmnet,service=wdqs-main
* 16:14 ryankemper: [[phab:T430880|T430880]] Stumbled across `wdqs2021` listed as inactive, looks like it was never fully re-pooled after a data xfer. Pooled.
* 16:12 btullis@deploy1003: Finished deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2] (duration: 00m 38s)
* 16:12 btullis@deploy1003: Started deploy [analytics/refinery@0e301af] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@0e301af2]
* 16:12 btullis@deploy1003: sync-world failed: <CalledProcessError> Command 'sudo -u mwbuilder /srv/mwbuilder/release/make-container-image/build-images.py --http-proxy http://webproxy:8080 --https-proxy http://webproxy:8080 /srv/mediawiki-staging/scap/image-build --staging-dir /srv/mediawiki-staging --mediawiki-versions 1.47.0-wmf.17,1.47.0-wmf.18,next --multiversion-image-basename docker-registry.discovery.wmnet/restricted/m
* 16:07 ryankemper@puppetserver1001: conftool action : set/pooled=yes; selector: name=wdqs101[1-4].eqiad.wmnet
* 16:03 btullis@deploy1003: Started scap sync-world: Rebuilding to pick up new version of dump scripts in mediawiki-cli for [[phab:T436913|T436913]]
* 16:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]] (duration: 09m 29s)
* 15:51 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=urldownloader[12]00[56].wikimedia.org [reason: depooling urldownloader trixie nodes]
* 15:51 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1325868{{!}}Growth: Remove now removed config variable]]
* 15:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:29 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:24 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:24 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2002.codfw.wmnet with reason: host reimage
* 15:18 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:15 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:13 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:09 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:05 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2002.codfw.wmnet with OS bookworm
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 15:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 15:04 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1073.eqiad.wmnet with reason: host reimage
* 15:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 15:02 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 15:00 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader
* 14:58 blake@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 14:53 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-ulsfo
* 14:49 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1073.eqiad.wmnet with OS trixie
* 14:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2002.codfw.wmnet
* 14:45 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1066.eqiad.wmnet with reason: host reimage
* 14:39 sukhe: sudo cumin "A:cp-text" "run-puppet-agent --enable 'merging CR 1334855'": [[phab:T425441|T425441]]
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:38 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2002.codfw.wmnet
* 14:32 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:30 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1066.eqiad.wmnet with OS trixie
* 14:29 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:29 sukhe: sudo cumin "A:cp-text" "disable-puppet 'merging CR 1334855'" [[phab:T425441|T425441]]
* 14:27 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-ulsfo
* 14:22 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2002.codfw.wmnet
* 14:21 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:19 arnaudb@dns1006: END - running authdns-update
* 14:18 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1074.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1074
* 14:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1074
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:17 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1074] - vriley@cumin1003"
* 14:17 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-codfw
* 14:17 arnaudb@dns1006: START - running authdns-update
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1073.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:16 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1072.eqiad.wmnet with reason: host reimage
* 14:16 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1073
* 14:15 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1073
* 14:12 ayounsi@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host netflow2004.codfw.wmnet with OS trixie
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1073] - vriley@cumin1003"
* 14:10 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:08 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 14:07 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] (duration: 09m 36s)
* 14:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1067.eqiad.wmnet with OS trixie
* 14:05 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:05 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 14:04 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 14:03 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:03 samtar@deploy1003: btullis, samtar: Continuing with deployment
* 14:02 samtar@deploy1003: btullis, samtar: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:01 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1072.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 14:01 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:58 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1328202{{!}}Remove the webrequest.dumps.dev0 stream from EventStreamConfig (T425087)]]
* 13:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:56 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - filippo@cumin1003"
* 13:55 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1054.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqiad and A:durum
* 13:52 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-codfw
* 13:51 moritzm: installing sqlite3 security updates
* 13:51 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:51 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqiad and A:durum
* 13:49 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-codfw and A:durum
* 13:48 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:47 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-codfw and A:durum
* 13:47 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-esams
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:44 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1067.eqiad.wmnet with reason: host reimage
* 13:43 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:43 ayounsi@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2004.codfw.wmnet with reason: host reimage
* 13:42 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] (duration: 13m 50s)
* 13:41 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1072.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:40 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1072
* 13:40 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1065.eqiad.wmnet with reason: host reimage
* 13:39 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-esams and A:durum
* 13:38 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1072
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:38 samtar@deploy1003: samtar, thiemowmde: Continuing with deployment
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:37 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1072] - vriley@cumin1003"
* 13:37 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-esams and A:durum
* 13:37 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-drmrs and A:durum
* 13:35 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:33 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:33 samtar@deploy1003: samtar, thiemowmde: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-eqsin and A:durum
* 13:31 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-eqsin and A:durum
* 13:28 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333866{{!}}Remove revisionId from term fallback cache lines for properties (T434204)]]
* 13:28 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1067.eqiad.wmnet with OS trixie
* 13:27 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:24 ayounsi@cumin1004: START - Cookbook sre.hosts.reimage for host netflow2004.codfw.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:22 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 13:22 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-esams
* 13:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:15 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:15 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:15 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1067.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 andrew@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 moritzm: installing bash updates from trixie point release
* 13:14 andrew@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cloudcephosd1045.eqiad.wmnet
* 13:14 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1066.eqiad.wmnet with OS trixie
* 13:14 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1067
* 13:13 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1067
* 13:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2901: Test
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:11 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:11 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1067] - vriley@cumin1003"
* 13:10 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 13:09 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:09 moritzm: installing libxslt bugfix updates from Trixie point release
* 13:08 jelto@dns1004: END - running authdns-update
* 13:06 jelto@dns1004: START - running authdns-update
* 13:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 13:05 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 13:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 13:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 13:04 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 13:00 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1066.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:59 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1066
* 12:59 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1066
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:58 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:58 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1066] - vriley@cumin1003"
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:56 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:56 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:54 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 12:53 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2901: Test
* 12:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2901: Test
* 12:52 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2901: Test
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:52 jayme@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'sync'.
* 12:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'sync'.
* 12:51 jayme@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'sync'.
* 12:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 12:50 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 12:50 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2901: Test
* 12:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2901: Test
* 12:50 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 12:49 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 12:45 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 12:42 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:41 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-ulsfo and A:durum
* 12:40 slyngshede@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling restart_daemons on A:durum-magru and A:durum
* 12:38 slyngshede@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling restart_daemons on A:durum-magru and A:durum
* 12:34 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 12:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1282: Pooling db1282 into s6
* 12:31 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:30 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:25 filippo@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 12:21 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:21 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 12:19 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 12:15 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_ulsfo
* 12:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1020.eqiad.wmnet with OS bookworm
* 12:10 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: db2207 repool
* 12:07 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_ulsfo
* 12:04 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 11:58 kart_: cxserver: Use urldownloader LVS endpoint ([[phab:T429175|T429175]])
* 11:57 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 11:56 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 11:56 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_magru
* 11:55 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 11:55 moritzm: installing rsync security updates
* 11:55 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 11:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:52 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 11:51 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 11:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1020.eqiad.wmnet with reason: host reimage
* 11:46 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1282: Pooling db1282 into s6
* 11:45 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1282 to dbctl [[phab:T407942|T407942]]', diff saved to https://phabricator.wikimedia.org/P96328 and previous config saved to /var/cache/conftool/dbconfig/20260903-114526-marostegui.json
* 11:43 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqiad
* 11:35 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqiad
* 11:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1020.eqiad.wmnet with OS bookworm
* 11:26 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 11:24 cgoubert@deploy1003: Finished scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter (duration: 12m 01s)
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: db2207 repool
* 11:22 cgoubert@deploy1003: cgoubert: Continuing with deployment
* 11:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_eqsin
* 11:17 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_esams
* 11:15 cgoubert@deploy1003: cgoubert: mediawiki: enable forward of fatal metrics to statsd exporter synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:14 cgoubert@deploy1003: Started scap sync-world: mediawiki: enable forward of fatal metrics to statsd exporter
* 11:10 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_esams
* 11:09 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 11:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1019.eqiad.wmnet with OS bookworm
* 11:01 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_drmrs
* 10:59 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-upload_codfw
* 10:52 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-upload_codfw
* 10:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:41 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 10:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1019.eqiad.wmnet with reason: host reimage
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow2005.codfw.wmnet
* 10:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2005.codfw.wmnet with OS trixie
* 10:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:17 btullis@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 10:17 btullis@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 10:16 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 10:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1019.eqiad.wmnet with OS bookworm
* 10:15 btullis@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-analytics: sync
* 10:15 btullis@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-analytics: sync
* 10:12 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s8
* 10:11 marostegui: Move s8 sanitarium from db1167 to db1281 [[phab:T434778|T434778]]
* 10:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 10:03 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow2005.codfw.wmnet with reason: host reimage
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:55 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow2003.codfw.wmnet with OS trixie
* 09:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:43 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow2005.codfw.wmnet with OS trixie
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:42 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow2005.codfw.wmnet on all recursors
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:41 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow2005.codfw.wmnet - ayounsi@cumin1003"
* 09:41 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1018.eqiad.wmnet with OS bookworm
* 09:41 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_ulsfo
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 09:39 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:39 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow2003.codfw.wmnet with reason: host reimage
* 09:39 hnowlan: fixed currently oncall pane in klaxon
* 09:38 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 09:38 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434751|T434751]]
* 09:37 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 09:37 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow2005.codfw.wmnet
* 09:35 marostegui: Move s7 sanitarium from db1158 to db1273 [[phab:T434775|T434775]]
* 09:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s7
* 09:34 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:33 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_ulsfo
* 09:33 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqiad
* 09:30 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] (duration: 09m 30s)
* 09:27 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 09:25 zabe@deploy1003: zabe: Continuing with deployment
* 09:25 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqiad
* 09:25 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 09:25 zabe@deploy1003: zabe: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1174 from dbctl [[phab:T436904|T436904]]', diff saved to https://phabricator.wikimedia.org/P96323 and previous config saved to /var/cache/conftool/dbconfig/20260903-092448-marostegui.json
* 09:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:21 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1334743{{!}}HookHandler: Fix bail out condition in onLocalUserCreated subscriber (T436791)]]
* 09:18 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_eqsin
* 09:17 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1018.eqiad.wmnet with reason: host reimage
* 09:15 topranks: put traffic on Lumen codfw<->eqiad link as it is stable [[phab:T435810|T435810]]
* 09:14 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_esams
* 09:09 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 09:06 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_esams
* 09:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 09:03 marostegui: Move s6 sanitarium from db1165 to db1279 [[phab:T434775|T434775]]
* 09:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host aqs1018.eqiad.wmnet with OS bookworm
* 08:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Changing sanitarium master in s6
* 08:57 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_drmrs
* 08:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_codfw
* 08:55 blake@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 08:49 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_codfw
* 08:49 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:45 slyngshede@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-purged (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 08:45 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cassandra-dev2001.codfw.wmnet with reason: host reimage
* 08:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:42 marostegui: Move s5 sanitarium from db1161 to db1275 [[phab:T434776|T434776]]
* 08:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 24 hosts with reason: Changing sanitarium master in s5
* 08:38 slyngshede@cumin1003: START - Cookbook sre.cdn.roll-restart-purged rolling restart_daemons on A:cp-text_magru
* 08:37 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:37 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server (duration: 01m 21s)
* 08:37 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:36 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to prod server
* 08:34 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1053.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:33 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server (duration: 01m 28s)
* 08:32 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@b146ae7] (releasing): [[phab:T436812|T436812]] to backup server
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cassandra-dev2001.codfw.wmnet with OS bookworm
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 08:11 moritzm: uploaded wmf-laptop 1.0.7 to apt.wikimedia.org
* 08:03 marostegui: Move s2 sanitarium from db1156 to db1271 [[phab:T434287|T434287]]
* 07:59 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:59 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cassandra-dev2001.codfw.wmnet
* 07:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 23 hosts with reason: Changing sanitarium master in s2
* 07:49 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host cassandra-dev2001.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cassandra-dev2001.codfw.wmnet
* 07:29 chlod: UTC morning backport window done
* 07:27 chlod@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] (duration: 11m 54s)
* 07:22 chlod@deploy1003: chlod, tryvix1509: Continuing with deployment
* 07:22 XioNoX: push pfw policies - [[phab:T436729|T436729]]
* 07:20 chlod@deploy1003: chlod, tryvix1509: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:15 chlod@deploy1003: Started scap sync-world: Backport for [[gerrit:1333872{{!}}core-Permissions.php: Allow English Wikiquote administrators to grant and remove the confirmed user permission (T436651)]]
* 07:15 marostegui: Power off db1228 for maintenance
* 07:13 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1228.eqiad.wmnet with reason: Onsite maintenance
* 07:01 arnaudb@dns1006: END - running authdns-update
* 06:58 arnaudb@dns1006: START - running authdns-update
* 06:54 jmm@cumin2003: END (PASS) - Cookbook sre.wdqs.restart-nginx-envoy (exit_code=0) rolling restart_daemons on A:wcqs-public
* 06:52 jmm@cumin2003: START - Cookbook sre.wdqs.restart-nginx-envoy rolling restart_daemons on A:wcqs-public
* 06:46 moritzm: installing libxml2 security updates
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 06:11 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 06:04 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 06:03 arnaudb@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts2002.codfw.wmnet
* 06:00 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 05:56 arnaudb@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts2002.codfw.wmnet
* 02:09 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 48s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:16 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1017.eqiad.wmnet with OS bookworm
== 2026-09-02 ==
* 23:57 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] (duration: 10m 21s)
* 23:51 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1017.eqiad.wmnet with reason: host reimage
* 23:50 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 23:49 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1334141{{!}}private/readme.php: Remove now removed secrets (T436880)]]
* 23:38 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:38 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:27 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 23:14 eevans@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host aqs1017.eqiad.wmnet with OS bookworm
* 22:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] (duration: 11m 03s)
* 22:32 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 22:30 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:26 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333307{{!}}Campaigns should not override existing campaign query strings (T436681)]]
* 22:05 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] (duration: 14m 14s)
* 21:59 krinkle@deploy1003: krinkle: Continuing with deployment
* 21:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:50 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1334068{{!}}Extract MathJax DOM filter into a separate file (T435274)]], [[gerrit:1334070{{!}}Respect contextual binomial sizing (T434477 T418144 T401718)]], [[gerrit:1334032{{!}}Remove the last vestiges of $wgVirtualRestConfig (T436054)]]
* 21:44 inflatador: bking@apt1002 sudo -E private_reprepro --ignore=wrongdistribution -C matomo_plugins include bookworm-wikimedia-private matomo-plugin-customreports_5.5.0-1_amd64.changes [[phab:T431608|T431608]]
* 21:40 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] (duration: 09m 48s)
* 21:35 jforrester@deploy1003: jforrester: Continuing with deployment
* 21:34 jforrester@deploy1003: jforrester: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:34 inflatador: bking@apt1002 sudo -E reprepro -C main include bookworm-wikimedia matomo-plugin-marketingcampaignsreporting_5.2.2-3_amd64.changes [[phab:T431608|T431608]]
* 21:30 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1324368{{!}}[WikiLambda] Log the …Orchestrator and …AbstractClient channels too]], [[gerrit:1333732{{!}}wikifunctions: Set up the functionmaintainer right for the community (T435637)]]
* 21:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:24 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1017.eqiad.wmnet with OS bookworm
* 21:19 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:19 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:11 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 21:04 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:03 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:01 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 20:50 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 20:27 dancy@deploy1003: Finished scap sync-world: testing (duration: 09m 21s)
* 20:18 dancy@deploy1003: Started scap sync-world: testing
* 20:18 dancy@deploy1003: Installation of scap version "4.288.0" completed for 3 hosts
* 20:16 dancy@deploy1003: Installing scap version "4.288.0" for 3 host(s)
* 19:57 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] (duration: 64m 27s)
* 19:55 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 18:57 jforrester@deploy1003: jforrester: Continuing with deployment
* 18:57 jforrester@deploy1003: jforrester: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 swfrench-wmf: deleted pods coredns-85b4f68d95-pk5sn coredns-85b4f68d95-22ddb coredns-85b4f68d95-49k5p in eqiad due to intermittent upstream resolution health check failures correlated with high DNS resolution latency
* 18:53 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1334006{{!}}Revert "Remove fallback to Special:Upload and redirect user to alternative form" (T436851)]]
* 18:31 sukhe@dns1004: END - running authdns-update
* 18:28 sukhe@dns1004: START - running authdns-update
* 18:26 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:26 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:17 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 18:16 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:16 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1065.eqiad.wmnet with OS trixie
* 18:16 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:14 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:14 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqiad
* 18:14 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 18:02 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:58 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:49 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqiad
* 17:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 17:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 17:45 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-eqsin
* 17:38 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 17:38 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 17:35 swfrench@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 17:20 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-eqsin
* 17:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 17:00 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1065.eqiad.wmnet with OS trixie
* 16:59 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:48 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 16:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1065.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 16:41 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1065
* 16:41 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1065
* 16:41 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_drmrs
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:38 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1065] - vriley@cumin1003"
* 16:34 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 16:29 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_drmrs
* 16:24 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_drmrs
* 16:14 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_drmrs
* 16:11 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-unlock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 16:10 root@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]] (duration: 48m 09s)
* 16:10 root@deploy1003: Forcefully removing global lock: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 16:10 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-unlock-scap for datacenter switchover from codfw to eqiad
* 16:10 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:59 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:58 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-run-puppet-on-db-masters for datacenter switchover from codfw to eqiad
* 15:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_eqsin
* 15:58 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.09-restore-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.09-restore-ttl for datacenter switchover from codfw to eqiad
* 15:57 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-start-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:57 root@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-cron: apply
* 15:57 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-cron: apply
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-start-maintenance for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 root@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-jobrunner: sync
* 15:56 root@deploy1003: helmfile [codfw] START helmfile.d/services/mw-jobrunner: sync
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.08-restart-mw-jobrunner for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.07-set-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:56 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period ends at: 2026-09-02 15:56:13.434320
* 15:56 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.07-set-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.06-set-db-readwrite for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.04-switch-mediawiki for datacenter switchover from codfw to eqiad
* 15:55 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.03-set-db-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.03-set-db-readonly for datacenter switchover from codfw to eqiad
* 15:54 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.02-set-readonly (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: [NON-PRIMARY-DC] MediaWiki read-only period starts at: 2026-09-02 15:53:47.690918
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.02-set-readonly for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.01-stop-maintenance (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.01-stop-maintenance for datacenter switchover from codfw to eqiad
* 15:53 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-reduce-ttl (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:47 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-reduce-ttl for datacenter switchover from codfw to eqiad
* 15:46 slyngshede@cumin1003: END (ERROR) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=97) for datacenter switchover from codfw to eqiad
* 15:45 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_eqsin
* 15:44 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_eqsin
* 15:42 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 15:39 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:38 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:31 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_eqsin
* 15:28 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:27 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-upload_magru
* 15:27 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin1004.eqiad.wmnet with reason: deploy homer to cumin1004 - cmooney@cumin1003
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from codfw to eqiad
* 15:22 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-lock-scap (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:22 root@deploy1003: Locking from deployment [ALL REPOSITORIES]: Datacenter switchover from codfw to eqiad - [[phab:T436781|T436781]]
* 15:22 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-lock-scap for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: END (PASS) - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks (exit_code=0) for datacenter switchover from codfw to eqiad
* 15:21 slyngshede@cumin1003: START - Cookbook sre.switchdc.mediawiki.00-downtime-db-readonly-checks for datacenter switchover from codfw to eqiad
* 15:17 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 15:15 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-upload_magru
* 15:11 moritzm: import jenkins 2.568.3 to thirdparty/jenkins for trixie-wikimedia
* 14:56 cdobbins@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-ats (exit_code=0) rolling restart_daemons on A:cp-text_magru
* 14:44 cdobbins@cumin1003: START - Cookbook sre.cdn.roll-restart-ats rolling restart_daemons on A:cp-text_magru
* 14:32 moritzm: installing pdns-recursor security updates
* 14:27 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:20 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:12 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:09 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:09 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:09 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:08 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:08 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:06 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:06 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:06 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:04 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 13:44 moritzm: bounce tcpircbot-logmsgbot/tcpircbot-logmsgbot_cloud on alert1002 to allow cumin1004 [[phab:T427897|T427897]]
* 13:36 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] (duration: 09m 52s)
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Continuing with deployment
* 13:31 samtar@deploy1003: jforrester, samtar, mlitn: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:26 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1333264{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333263{{!}}Add thumb.* to allowed hosts for commons images (T436579)]], [[gerrit:1333816{{!}}Enable mobile MMV on all wikis (T429970)]]
* 13:25 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 13:24 moritzm: installing wireshark security updates
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:24 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 13:23 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 13:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:04 moritzm: import librsvg 2.60.0+dfsg-1+wmf13u1 to component/thumbor for trixie-wikimedia [[phab:T436505|T436505]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-experiment-platform: apply
* 12:44 atsuko@dns1004: END - running authdns-update
* 12:41 atsuko@dns1004: START - running authdns-update
* 12:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:35 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/push-notifications: apply
* 12:34 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/push-notifications: apply
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] (duration: 12m 50s)
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:25 jelto@cumin1003: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:24 dreamyjazz@deploy1003: dreamyjazz, btullis: Continuing with deployment
* 12:24 jelto@cumin1003: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:22 dreamyjazz@deploy1003: dreamyjazz, btullis: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 jelto@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:17 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1328201{{!}}Declare the webrequest.dumps.v1 stream in EventStreamConfig (T425087 T291645)]], [[gerrit:1329360{{!}}EventStreamConfig: Register the abuse_review_interaction stream (T435517)]]
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'liftwing-openapi-server' for release 'main' .
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 11:31 marostegui@cumin1003: Removing db1172 from zarcillo [[phab:T436763|T436763]]
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1172.eqiad.wmnet
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:31 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:30 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1172.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 11:26 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:26 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:25 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 11:25 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:24 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:23 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:20 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1172.eqiad.wmnet
* 11:20 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 11:12 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:11 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:09 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:08 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:05 slyngshede@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.*
* 11:05 slyngshede@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 11:05 slyngshede@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 11:03 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:03 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:00 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 10:59 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 10:52 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:50 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:48 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Pool back db1242
* 10:45 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow4003.ulsfo.wmnet with OS trixie
* 10:31 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1172 from dbctl [[phab:T436763|T436763]]', diff saved to https://phabricator.wikimedia.org/P96318 and previous config saved to /var/cache/conftool/dbconfig/20260902-103152-marostegui.json
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow4003.ulsfo.wmnet with reason: host reimage
* 10:08 blake@deploy1003: Finished scap sync-world: non-build deployment for [[phab:T417800|T417800]] (duration: 05m 37s)
* 10:06 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:05 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:04 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:03 blake@deploy1003: Started scap sync-world: non-build deployment for [[phab:T417800|T417800]]
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 10:00 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 10:00 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:59 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db1242: Pool back db1242
* 09:58 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:57 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:56 jmm@dns1004: END - running authdns-update
* 09:56 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:55 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:54 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:53 jmm@dns1004: START - running authdns-update
* 09:51 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Pool back db1242
* 09:47 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1228 to dbctl [[phab:T435892|T435892]]', diff saved to https://phabricator.wikimedia.org/P96313 and previous config saved to /var/cache/conftool/dbconfig/20260902-094713-marostegui.json
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:44 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:42 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow4003.ulsfo.wmnet with OS trixie
* 09:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow7002.magru.wmnet with OS trixie
* 09:31 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:30 moritzm: installing openjdk-21 security updates
* 09:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:22 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 09:17 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:17 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 09:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 09:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 09:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:10 moritzm: installing openjdk-8 security updates
* 09:09 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow7002.magru.wmnet with reason: host reimage
* 09:08 elukey@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 08:46 tappof: bump space for prometheus k8s-dse in eqiad
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:39 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow7002.magru.wmnet with OS trixie
* 08:36 moritzm: installing libgraphite2 security updates
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 08:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 08:23 Msz2001: UTC morning backport window done
* 08:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] (duration: 14m 36s)
* 08:19 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:19 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 08:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 08:14 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 08:14 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 08:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 08:08 fabfur@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cp5022.eqsin.wmnet with reason: investigating
* 08:07 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1332684{{!}}Update stream config for user_info_card_interaction (T435585)]]
* 08:07 slyngshede@cumin1003: conftool action : set/pooled=no; selector: name=cp5022.*
* 08:07 fabfur: depooling and silencing cp5022 ([[phab:T414411|T414411]])
* 08:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 08:03 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:03 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* {{safesubst:SAL entry|1=08:03 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)}}
* 07:49 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 07:49 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]] synced to the
* 07:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1004.eqiad.wmnet
* 07:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1004.eqiad.wmnet
* 07:30 jmm@dns1004: END - running authdns-update
* {{safesubst:SAL entry|1=07:27 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333559{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333560{{!}}UserInfoCard: Send the source page with the api_request event (T435585)]], [[gerrit:1333561{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]], [[gerrit:1333562{{!}}UserInfoCard: Send the place of the trigger with api_request (T435585)]}}
* 07:27 jmm@dns1004: START - running authdns-update
* 07:22 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] (duration: 16m 04s)
* 07:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1242: Cloning db1228
* 07:20 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1242: Cloning db1228
* 07:18 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1228,1242].eqiad.wmnet with reason: db1242 needs to clone db1228
* 07:17 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Continuing with deployment
* 07:12 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1228.eqiad.wmnet with OS trixie
* 07:10 mszwarc@deploy1003: mszwarc, eggroll97, tryvix1509: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be veri
* 07:06 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1333338{{!}}throttle.php: Lift IP cap for Mapudungun editathon on 2026-09-05 (T436672)]], [[gerrit:1333343{{!}}InitialiseSettings.php: Set $wgUploadNavigationUrl for ukwiki (T436712)]], [[gerrit:1333390{{!}}Add two wmf groups to privileged status (T436734)]]
* 06:43 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: build: Updating npm dependencies (duration: 00m 13s)
* 06:43 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: build: Updating npm dependencies
* 06:39 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:32 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1228.eqiad.wmnet with reason: host reimage
* 06:23 slyngshede@dns1004: END - running authdns-update
* 06:21 marostegui: Drop cu* tables from s3 bswiktionary [[phab:T435965|T435965]]
* 06:20 slyngshede@dns1004: START - running authdns-update
* 06:18 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1228.eqiad.wmnet with OS trixie
* 06:13 XioNoX: re-enable magru cr1/asw1-b3 link - [[phab:T436675|T436675]]
* 05:06 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] (duration: 04m 42s)
* 05:04 tstarling@deploy1003: tstarling: Continuing with deployment
* 05:03 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 05:01 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 05:01 tstarling@deploy1003: Scap cancelled without rolling back.
* 04:53 tstarling@deploy1003: tstarling: Continuing with deployment
* 04:29 tstarling@deploy1003: tstarling: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:25 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1333470{{!}}Updater: Normalize MW_VERSION (T436741)]], [[gerrit:1333423{{!}}Set a short CC:max-age on cacheable REST responses]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 43s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-01 ==
* 21:59 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] (duration: 18m 05s)
* 21:52 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:47 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:41 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333294{{!}}Merge branch 'master' into wmf_deploy]]
* 21:38 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] (duration: 23m 55s)
* 21:28 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 21:20 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:14 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1333293{{!}}Merge branch 'master' into wmf_deploy]], [[gerrit:1333277{{!}}Preserve showlogin query parameter on redirects (T435248)]]
* 20:47 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aqs1016.eqiad.wmnet with OS bookworm
* 20:28 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:24 eevans@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aqs1016.eqiad.wmnet with reason: host reimage
* 20:11 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 20:11 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:57 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:57 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:56 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:56 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:54 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:52 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:47 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:45 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:42 eevans@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:41 jhancock@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:40 eevans@cumin1003: START - Cookbook sre.hosts.provision for host aqs1016.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:39 eevans@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host aqs1016.eqiad.wmnet with OS bookworm
* 19:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:32 jhancock@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudcephosd1055
* 19:32 jhancock@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:30 cdobbins@cumin1003: conftool action : set/pooled=yes; selector: name=cp5022.* [reason: update IP addrs]
* 19:30 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for cp5022.eqsin.wmnet
* 19:30 cdobbins@cumin1003: START - Cookbook sre.hosts.remove-downtime for cp5022.eqsin.wmnet
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:25 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:24 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:23 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:22 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:21 eevans@cumin1003: START - Cookbook sre.hosts.reimage for host aqs1016.eqiad.wmnet with OS bookworm
* 19:16 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1056.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:14 jclark@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:13 jclark@cumin1003: START - Cookbook sre.hosts.provision for host cloudcephosd1055.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1056
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1056
* 19:12 jclark@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudcephosd1055
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1055
* 19:12 jclark@cumin1003: END (ERROR) - Cookbook sre.network.configure-switch-interfaces (exit_code=97) for host cloudcephosd1054
* 19:12 jclark@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudcephosd1054
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:11 jclark@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:11 jclark@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding cloudcephosd1055 to eqiad - jclark@cumin1003"
* 19:06 jclark@cumin1003: START - Cookbook sre.dns.netbox
* 19:06 sukhe@dns1004: END - running authdns-update
* 19:03 sukhe@dns1004: START - running authdns-update
* 18:14 dancy@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 17:04 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on matomo1003.eqiad.wmnet with reason: Matomo version upgrade [[phab:T431608|T431608]]
* 17:00 dancy@deploy1003: Finished scap sync-world: testing (duration: 08m 07s)
* 16:52 dancy@deploy1003: Started scap sync-world: testing
* 16:48 dancy@deploy1003: sync-world aborted: testing (duration: 00m 05s)
* 16:48 dancy@deploy1003: Started scap sync-world: testing
* 16:47 dancy@deploy1003: Installation of scap version "4.287.0" completed for 156 hosts
* 16:42 dancy@deploy1003: Installing scap version "4.287.0" for 156 host(s)
* 16:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 16:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 15:51 moritzm: installing mesa security updates
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts phab1004.eqiad.wmnet
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:29 aokoth@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:27 aokoth@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: phab1004.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - aokoth@cumin1003"
* 15:20 aokoth@cumin1003: START - Cookbook sre.dns.netbox
* 15:14 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595] (duration: 05m 32s)
* 15:14 aokoth@cumin1003: START - Cookbook sre.hosts.decommission for hosts phab1004.eqiad.wmnet
* 15:11 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on phab1004.eqiad.wmnet with reason: Decom
* 15:09 joal@deploy1003: Started deploy [analytics/refinery@19f4b59]: Regular analytics weekly train [analytics/refinery@19f4b595]
* 15:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config: apply
* 15:08 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config: apply
* 14:55 hashar: Restarted Jenkins on releases1003
* 14:51 hashar: Restarted CI Jenkins on contint1003
* 14:48 hashar: Restarting Gerrit primary on gerrit2003
* 14:45 hashar: Restarted Gerrit on gerrit1003 and gerrit2002
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 14:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 14:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 14:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:27 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-content-change-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:25 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich: apply
* 14:24 moritzm: installing curl security updates
* 14:24 jmm@dns1004: END - running authdns-update
* 14:23 hashar@deploy1003: Finished deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser (duration: 00m 14s)
* 14:23 hashar@deploy1003: Started deploy [integration/docroot@780c44c]: opensource.yaml: Add CheckUser
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 14:21 jmm@dns1004: START - running authdns-update
* 14:21 jmm@dns1004: END - running authdns-update
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 14:20 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 14:19 jmm@dns1004: START - running authdns-update
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 14:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 14:14 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:14 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:13 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595] (duration: 07m 26s)
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Test
* 14:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:10 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Test
* 14:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 14:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Test
* 14:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Test
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 14:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 14:05 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (thin): Regular analytics weekly train THIN [analytics/refinery@19f4b595]
* 14:05 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:04 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] (duration: 37m 37s)
* 14:04 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:03 hashar: Removed openjdk-17 packages from contint1002/contint2002 following relocation of CI Jenkins to contint1003/contint2003 # [[phab:T418521|T418521]]
* 14:03 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 14:03 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:02 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 14:02 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 14:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:00 joal@deploy1003: Finished deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595] (duration: 00m 45s)
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 13:59 joal@deploy1003: Started deploy [analytics/refinery@19f4b59] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@19f4b595]
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:58 ladsgroup@dns1004: END - running authdns-update
* 13:58 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 13:57 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 ladsgroup@dns1004: START - running authdns-update
* 13:56 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:56 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:56 ladsgroup@dns1004: END - running authdns-update
* 13:55 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:53 ladsgroup@dns1004: START - running authdns-update
* 13:50 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 13:49 kharlan@deploy1003: kharlan: Continuing with deployment
* 13:48 kharlan@deploy1003: kharlan: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:43 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2004.wikimedia.org
* 13:42 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 13:41 jmm@dns1004: END - running authdns-update
* 13:40 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2004.wikimedia.org
* 13:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2003.wikimedia.org
* 13:38 jmm@dns1004: START - running authdns-update
* 13:34 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2003.wikimedia.org
* 13:33 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:29 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1004.wikimedia.org
* 13:26 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1333200{{!}}Special:AbuseReview: Show changes as core's inline diff (T436490)]]
* 13:26 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:25 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:24 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1004.wikimedia.org
* 13:24 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 13:23 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] (duration: 20m 24s)
* 13:20 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1003.wikimedia.org
* 13:20 fnegri@deploy1003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 13:19 fnegri@deploy1003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 13:16 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1003.wikimedia.org
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:16 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 13:15 moritzm: bump urldownloader[12]00[34] to 8G RAM [[phab:T429175|T429175]]
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 13:15 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 13:15 fnegri@deploy1003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 13:14 fnegri@deploy1003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:12 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 13:11 fnegri@deploy1003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 13:11 aude@deploy1003: aude: Continuing with deployment
* 13:10 fnegri@deploy1003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 13:07 aude@deploy1003: aude: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich-next: apply
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1332717{{!}}Enable ReadingLists for logged-in users on phase 1 wikis (T434922)]]
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:02 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-content-history-reconcile-enrich-next: apply
* 13:01 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/datasets-config-next: apply
* 13:00 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:59 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2004.wikimedia.org with OS bookworm
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host netflow1004.eqiad.wmnet
* 12:52 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host netflow1004.eqiad.wmnet with OS trixie
* 12:48 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:47 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 12:41 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] (duration: 16m 25s)
* 12:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:38 ayounsi@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:36 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 12:35 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 12:34 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:34 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 12:33 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2004.wikimedia.org with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 ayounsi@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on netflow1004.eqiad.wmnet with reason: host reimage
* 12:32 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:32 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 12:31 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:31 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:30 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 12:29 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:28 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 12:25 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1333182{{!}}EventStreamConfig: Fix User-Agent stream config for some schemas (T432848)]]
* 12:22 ayounsi@cumin1003: START - Cookbook sre.hosts.reimage for host netflow1004.eqiad.wmnet with OS trixie
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache netflow1004.eqiad.wmnet on all recursors
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:21 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:21 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM netflow1004.eqiad.wmnet - ayounsi@cumin1003"
* 12:20 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 12:20 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 12:16 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 12:16 ayounsi@cumin1003: START - Cookbook sre.ganeti.makevm for new host netflow1004.eqiad.wmnet
* 12:15 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2004.wikimedia.org with OS bookworm
* 12:14 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:09 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:07 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ipoid: apply
* 12:05 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader2003.wikimedia.org with OS bookworm
* 11:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:43 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader2003.wikimedia.org with reason: host reimage
* 11:32 jmm@cumin2003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling restart_daemons on A:kafka-test-eqiad
* 11:26 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader2003.wikimedia.org with OS bookworm
* 11:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1004.wikimedia.org with OS bookworm
* 11:12 moritzm: installing openjdk-21 security updates
* 11:12 jmm@cumin2003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling restart_daemons on A:kafka-test-eqiad
* 10:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:53 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1004.wikimedia.org with reason: host reimage
* 10:49 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2902: Pool back db2902
* 10:45 moritzm: installing Python 3.11 security updates
* 10:37 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1004.wikimedia.org with OS bookworm
* 10:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp5022.eqsin.wmnet with OS trixie
* 10:36 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp5022.eqsin.wmnet on all recursors
* 10:36 ayounsi@cumin1003: START - Cookbook sre.dns.wipe-cache cp5022.eqsin.wmnet on all recursors
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:34 ayounsi@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:34 ayounsi@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cp5022 - ayounsi@cumin1003"
* 10:30 ayounsi@cumin1003: START - Cookbook sre.dns.netbox
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:20 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:06 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:04 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:04 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2902: Pool back db2902
* 10:04 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2902: Pool back db2902
* 10:03 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:03 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2902: test
* 10:03 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 10:01 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 10:01 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:58 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool db2902: test
* 09:58 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 09:50 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/echoserver: apply
* 09:49 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/echoserver: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:24 marostegui@cumin1003: dbctl commit (dc=all): 'Change db1176 and db2230's weight, test-s4 masters, to 0 to mimic the rest of production [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96292 and previous config saved to /var/cache/conftool/dbconfig/20260901-092444-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96291 and previous config saved to /var/cache/conftool/dbconfig/20260901-090233-marostegui.json
* 09:02 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96290 and previous config saved to /var/cache/conftool/dbconfig/20260901-090158-marostegui.json
* 09:01 marostegui@cumin1003: dbctl commit (dc=all): 'Add db1901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96289 and previous config saved to /var/cache/conftool/dbconfig/20260901-090121-marostegui.json
* 09:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:56 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cp5022.eqsin.wmnet with reason: host reimage
* 08:50 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host urldownloader1003.wikimedia.org with OS bookworm
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1074.eqiad.wmnet
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:46 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1074.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:43 marostegui@cumin1003: dbctl commit (dc=all): 'Test repool db2902', diff saved to https://phabricator.wikimedia.org/P96288 and previous config saved to /var/cache/conftool/dbconfig/20260901-084317-marostegui.json
* 08:42 marostegui@cumin1003: dbctl commit (dc=all): 'Test depool db2902', diff saved to https://phabricator.wikimedia.org/P96287 and previous config saved to /var/cache/conftool/dbconfig/20260901-084249-marostegui.json
* 08:39 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:37 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2902: test
* 08:36 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db2902: test
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2903 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96285 and previous config saved to /var/cache/conftool/dbconfig/20260901-083557-marostegui.json
* 08:35 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2902 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96284 and previous config saved to /var/cache/conftool/dbconfig/20260901-083527-marostegui.json
* 08:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:34 marostegui@cumin1003: dbctl commit (dc=all): 'Add db2901 to test-s4 [[phab:T427059|T427059]]', diff saved to https://phabricator.wikimedia.org/P96283 and previous config saved to /var/cache/conftool/dbconfig/20260901-083432-marostegui.json
* 08:32 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1074.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1073.eqiad.wmnet
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:27 jmm@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on urldownloader1003.wikimedia.org with reason: host reimage
* 08:25 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1073.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 08:21 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['cp5022']
* 08:16 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1073.eqiad.wmnet
* 08:15 jmm@cumin2003: START - Cookbook sre.hosts.reimage for host urldownloader1003.wikimedia.org with OS bookworm
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1072.eqiad.wmnet
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:15 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:14 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1072.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:10 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 08:05 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1072.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1067.eqiad.wmnet
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:02 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 08:00 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1067.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:56 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:54 elukey@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022']
* 07:53 joal@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: apply
* 07:52 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 07:52 joal@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: apply
* 07:50 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1067.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1066.eqiad.wmnet
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:47 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:47 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1066.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 07:41 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:34 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1066.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1065.eqiad.wmnet
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:32 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:32 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1065.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:27 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 07:23 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1065.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1075.eqiad.wmnet
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:17 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 07:15 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1075.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 06:49 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 06:45 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1075.eqiad.wmnet
* 06:29 moritzm: installing Java 17 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.15 (duration: 02m 25s)
* 03:50 denisse@deploy1003: Finished deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2 (duration: 00m 19s)
* 03:50 denisse@deploy1003: Started deploy [librenms/librenms@9b0b502]: Upgrade LibreNMS to 26.8.2
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]] (duration: 37m 30s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.18 refs [[phab:T430837|T430837]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 33s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:30 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 00:29 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 00:21 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 00:18 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
0l7vy3qbyt5kmcln8ybglbrh9p8jlo1
Release Engineering/SAL
0
17290
2456785
2456713
2026-09-11T16:29:01Z
Stashbot
7414
dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1339162
2456785
wikitext
text/x-wiki
=== 2026-09-11 ===
* 16:29 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1339162
* 06:05 phedenskog: Updating Jenkins jobs from https://gerrit.wikimedia.org/r/c/integration/config/+/1338086 (quibble*, api-testing-*, mwext-phpunit-coverage*): run npm cache verify at most once a day per cache [[phab:T437377|T437377]]
=== 2026-09-10 ===
* 22:01 Southparkfan: decommission deployment-mwlog02 - [[phab:T436469|T436469]]
* 21:08 Southparkfan: masked and stopped stray udp2log.service on deployment-mwlog03, claimed 8420/udp which should have been assigned to udp_tee - [[phab:T436469|T436469]]
* 18:56 James_F: Zuul: [mediawiki/extensions/PersonalDashboard] Add ORES & Wikibase phan deps for [[phab:T436570|T436570]] and [[phab:T437491|T437491]]
* 18:55 James_F: Zuul: [mediawiki/extensions/ArticleGuidance] Add CommunityConfiguration for phan for [[phab:T437586|T437586]]
* 05:49 phedenskog: Updated quibble-for-mediawiki-core postgres and sqlite jobs to run only the phpunit-database stage [[phab:T437134|T437134]]
* 05:20 phedenskog: Updated 134 Quibble Jenkins jobs to drop the duplicated --reporting-url [[phab:T323750|T323750]]
=== 2026-09-09 ===
* 22:36 Southparkfan: point profile::rsyslog::udp_tee::destinations to both deployment-mwlog02 and deployment-mwlog03 8421/udp - [[phab:T436469|T436469]]
* 22:32 Southparkfan: set role::logging::mediawiki::udp2log::monitor: false to unbreak [[phab:T436469|T436469]]
* 16:53 Southparkfan: cpjobqueue: switch back to deployment-jobrunner05 (PHP 8.3) - [[phab:T435393|T435393]]
* 16:32 Southparkfan: adjust cpjobqueue config to temporarily send jobs to deployment-jobrunner06 (PHP 8.5) - [[phab:T435393|T435393]]
* 16:12 Southparkfan: decommission deployment-docker-mathoid02 - [[phab:T436465|T436465]]
* 16:09 thcipriani: deployed some anti-scraper mitigations on beta
* 15:40 komla@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.quota_increase (exit_code=0) by 30 server-groups ([[phab:T437117|T437117]])
* 15:40 komla@cloudcumin1001: START - Cookbook wmcs.openstack.quota_increase by 30 server-groups ([[phab:T437117|T437117]])
* 11:58 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.add_user_to_project (exit_code=0) for user 'denisse' in role 'member'
* 11:58 fnegri@cloudcumin1001: START - Cookbook wmcs.vps.add_user_to_project for user 'denisse' in role 'member'
* 09:42 phedenskog: jjb: update quibble-with-gated-extensions-selenium-php83 to run npm install ahead of the browser tests [[phab:T436376|T436376]]
=== 2026-09-08 ===
* 19:36 Southparkfan: truncate Apache2 forensic log on deployment-mediawiki14, disk almost full
* 19:35 Southparkfan: set profile::mediawiki::httpd::enable_forensic_log to false, to avoid disk exhaustion during high-traffic load
* 16:21 thcipriani: hard reboot deployment-prep:deployment-cache-text08 horizon logs showing OOMs
* 15:42 hashar: Updating Quibble jobs to 1.21.0 # [[phab:T429715|T429715]] [[phab:T303270|T303270]] [[phab:T437134|T437134]] [[phab:T300727|T300727]]
* 15:06 hashar: Tag Quibble @ {{Gerrit|b692707e5213abf5e16894b1fb21cea41cd72c9e}} # [[phab:T429715|T429715]] [[phab:T303270|T303270]] [[phab:T437134|T437134]] [[phab:T300727|T300727]]
=== 2026-09-07 ===
* 19:31 Southparkfan: 19:31 UTC: switch back from PHP 8.5 to PHP 8.3 hosts - [[phab:T435393|T435393]]
* 19:06 Southparkfan: 18:05 UTC: pool mediawiki15 and mediawiki16 (PHP 8.5) as replacements for 13 and 14 (PHP 8.3), smoke test - [[phab:T435393|T435393]]
* 19:03 Southparkfan: banhammer lots of ranges to get Beta Cluster back online; not sure it was very effective, but we seem to be out of the woods
* 17:55 Southparkfan: no disk space left on deployment-mediawiki14, cleared logs in /var/log/apache2/forensic to unbreak
=== 2026-09-04 ===
* 15:52 hashar: integration: remove from Jenkins global config: NPM_CONFIG_AUDIT=false and NPM_CONFIG_FUND=false # [[phab:T437008|T437008]]
* 14:43 hashar: integration: set in Jenkins global config: NPM_CONFIG_AUDIT=false and NPM_CONFIG_FUND=false # [[phab:T437008|T437008]]
=== 2026-09-03 ===
* 23:29 Southparkfan: attach Cinder volume for /srv on deploy06, jobrunner06, mediawiki15/mediawiki16 - [[phab:T435393|T435393]]
* 11:26 hashar: Deleted coverage report for SimilarEditors ( /srv/doc/cover-extensions/SimilarEditors ), extension is being archived # [[phab:T436880|T436880]]
* 08:13 James_F: Zuul: [mediawiki/services/similar-users] Archive service, for [[phab:T368269|T368269]]
* 07:25 hashar: integration: granted sudo access to Phedenskog
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 05:15 hashar: Updating tox jobs to change default python from 3.9 to 3.11 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1334023 {{!}} [[phab:T436857|T436857]]
=== 2026-09-02 ===
* 22:21 Southparkfan: decommission deployment-webperf21 - [[phab:T436464|T436464]]
* 17:09 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1333891
* 15:37 hashar: jjb: update Quibble jobs to 1.20.0 # [[phab:T321617|T321617]] [[phab:T432966|T432966]] [[phab:T428642|T428642]] [[phab:T435975|T435975]]
* 15:14 hashar: Building Quibble 1.20.0 images
* 14:52 hashar: Tag Quibble 1.20.0 @ {{Gerrit|c6f13962c27685773fc1fb39d74189b202e0961d}} # [[phab:T321617|T321617]] [[phab:T432966|T432966]] [[phab:T428642|T428642]] [[phab:T435975|T435975]]
=== 2026-09-01 ===
* 21:45 Southparkfan: hiera: switch ATS routing for performance.beta.wmcloud.org from webperf21 to webperf31 - [[phab:T436464|T436464]]
* 20:55 Southparkfan: decommission deployment-webperf22 - [[phab:T436464|T436464]]
* 20:52 Southparkfan: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1333288 cherry-picked on project puppetserver - [[phab:T436464|T436464]]
* 19:59 Southparkfan: decommission deployment-webperf32 - [[phab:T436464|T436464]]
* 17:54 Southparkfan: switch webproxy for wikifeeds-beta to deployment-docker-wikifeeds01 - [[phab:T436462|T436462]]
* 17:47 Southparkfan: switch profile::restbase::citoid_uri and profile::restbase::cxserver_uri to resp. citoid03 and cxserver03, old VMs no longer exist - [[phab:T436619|T436619]]
* 17:43 Southparkfan: revoked Puppet certs for deployment-docker-cxserver02 and deployment-docker-citoid02 - [[phab:T436619|T436619]]
* 14:36 hashar: integration: deleted Cypress from codehealth job after the job learn to instruct Cypress & Puppeter to no more download binary blobs. `rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-codehealth-master-non-voting/Cypress` # [[phab:T427471|T427471]]
=== 2026-08-31 ===
* 22:17 andrewbogott: (log again, mentioned wrong task last time) add PHP 8.5 hosts to Scap dsh groups - [[phab:T435393|T435393]] (andrew retrying SPF's failed log)
* 21:21 Southparkfan: switched cache-text08 backend from mediawiki14 to mediawiki16, then switched back to mediawiki14 to match Puppet state - [[phab:T435393|T435393]]
* 21:11 Southparkfan: add PHP 8.5 hosts to Scap dsh groups - [[phab:T436462|T436462]]
* 18:38 Southparkfan: decommission deployment-wikifeeds02 - [[phab:T436462|T436462]]
* 13:21 jnuche: Updating development images on contint primary for [[phab:T435931|T435931]]
* 09:16 elukey: move cx-server and citoid-beta endpoints in deployment-prep to two new Trixie VMs, update their configs and Docker images and delete the old images.
* 04:18 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1332530
=== 2026-08-27 ===
* 15:06 Krinkle: Disable duplicate publishing noise from extension-IPReputation, EIPR, [[phab:T143162|T143162]]
=== 2026-08-26 ===
* 18:07 andrewbogott: resolving rebase conflicts in /srv/git/labs/private
* 17:24 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1329617
* 14:45 hashar: Updating Castor (0.4.3..0.4.5) on all Jenkins jobs {{!}} https://gerrit.wikimedia.org/r/1329577
* 07:11 hashar: integration: updated Quibble jobs to replace deprecated `--commands` option by 1..n `--command` option(s) {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1293711 {{!}} [[phab:T321617|T321617]]
=== 2026-08-25 ===
* 21:32 brett: Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08
* 21:23 brett: revert deployment-prep acme-chief switch: acme-chief active to deployment-acme-chief05 and passive to deployment-acme-chief06
* 21:03 brett: Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08
* 20:31 Southparkfan: [[phab:T401839|T401839]] - provisioned deployment-docker-wikifeeds01 (trixie) using Tofu (cloudvps-repos/deployment-prep/tofu-provisioning)
* 15:06 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/121 ([[phab:T435368|T435368]])
* 07:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.remove_instance (exit_code=0) for instance deployment-ircd03
* 07:35 taavi@cloudcumin1001: START - Cookbook wmcs.vps.remove_instance for instance deployment-ircd03
=== 2026-08-24 ===
* 08:29 hashar: integration: add firewall rule to ssh from contint1003/contint2003 IPv6 (IPv4 was already allowed) # [[phab:T435756|T435756]]
* 08:26 hashar: integration: removing firewall rule for ssh from contint1002/contint2002 IPv4 # [[phab:T418521|T418521]]
=== 2026-08-21 ===
* 13:07 hashar: integration: triggered Doxygen doc for PersonalDashboard extension using: `zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/extensions/PersonalDashboard --change {{Gerrit|1328044}},1` # [[phab:T435392|T435392]]
=== 2026-08-20 ===
* 23:26 bd808: deployment-mediawiki14: `systemctl stop php8.3-fpm; sleep 5m; systemctl start php8.3-fpm` -- maybe the bot storm will break if we give fast 500 responses for 5 minutes.
* 14:55 James_F: jforrester@integration-castor06:$ sudo rm -rf /srv/castor/castor-mw-ext-and-skins/master/wikilambda-catalyst-end-to-end # Clear stale Catalyst castor npm downloads.
* 13:35 James_F: Zuul: [mediawiki/extensions/WikiLambda] Re-enable Catalyst
* 08:54 hashar: deployment-prep: hard reboot deployment-cache-test-08 # [[phab:T435421|T435421]]
=== 2026-08-19 ===
* 14:48 hashar: integration: updated castor save job to have rsync emit statistics {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1321049 {{!}} [[phab:T432685|T432685]]
* 12:30 James_F: Updating development images on contint primary for “fundraising: Drop libc-client-dev from the bookworm PHP 8.2 image”
=== 2026-08-18 ===
* 17:32 dancy: Rebooting deployment-mediawiki14.deployment-prep.eqiad1.wikimedia.cloud for good measure
* 17:31 dancy: rm /var/log/apache2/*.gz on deployment-mediawiki14.deployment-prep.eqiad1.wikimedia.cloud to free up ~6GB.
* 09:24 hashar: zuul: restarted zuul-web
=== 2026-08-17 ===
* 21:09 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/115
* 21:01 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/118 ([[phab:T413817|T413817]], [[phab:T407430|T407430]])
* 18:08 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/113
* 16:30 James_F: Zuul: [mediawiki/extensions/WP25EasterEggs] Archive repository, for [[phab:T418134|T418134]]
=== 2026-08-14 ===
* 18:59 James_F: Zuul: Enforce CI for mediawiki-php-<nowiki>{</nowiki>excimer,luasandbox,wikidiff2<nowiki>}</nowiki> for [[phab:T425943|T425943]]
* 18:20 James_F: Docker: [php85] Migrate to Wikimedia-provide binary, cascaded, for [[phab:T433254|T433254]]
* 14:37 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1325900
=== 2026-08-13 ===
* 19:55 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/111 ([[phab:T401115|T401115]]) (redux for typo fix)
* 19:27 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/111 ([[phab:T401115|T401115]])
* 03:05 TimStarling: created Produnto tables on beta [[phab:T421436|T421436]]
=== 2026-08-12 ===
* 12:31 hashar: integration: on Castor: `sudo rm -fR /srv/castor/*/*/mwext-phpunit-coverage*/npm` # [[phab:T427922|T427922]]
* 11:59 hashar: integration: on Castor: `sudo rm -fR /srv/castor/*/*/*codehealth*/npm` # [[phab:T427822|T427822]]
=== 2026-08-07 ===
* 16:13 hashar: integration: deleted integration-agent-[[phab:T422258|T422258]] agent # [[phab:T422258|T422258]]
* 14:31 hashar: integration: updating Quibble jobs to Quibble 1.19.0 # [[phab:T432934|T432934]] [[phab:T432943|T432943]] [[phab:T427922|T427922]]
* 07:25 hashar: Tag Quibble 1.19.0 @ {{Gerrit|a8a84ed1c0adb34688ce22066ebde8a2d584a7d4}} # [[phab:T432934|T432934]] [[phab:T432943|T432943]] [[phab:T427922|T427922]]
=== 2026-08-06 ===
* 18:02 dancy: Restarted gitlab-webhooks ([[phab:T430410|T430410]]) (revert)
* 17:55 dancy: Restarted gitlab-webhooks ([[phab:T430410|T430410]])
=== 2026-08-05 ===
* 20:31 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1184176
* 18:39 James_F: Docker: [php85] Upgrade PHP to 8.5.9
* 15:58 James_F: Zuul: [mediawiki/extensions/WikiLambda] Disable Catalyst, corrupted npm cache
=== 2026-08-04 ===
* 19:06 dancy: Updating buildkitd to v0.32.2 on gitlab-cloud-runners (staging and production) ([[phab:T433985|T433985]])
* 16:48 James_F: Docker: [php83] Upgrade PHP to 8.3.33
=== 2026-08-03 ===
* 20:45 dancy: Updating buildkitd to v0.32.1 on gitlab-cloud-runners (staging and production) ([[phab:T433879|T433879]])
* 16:40 hashar: gerrit: added Vaughn Walters to integration group until he get added to the ciadmin LDAP group {{!}} [[phab:T433615|T433615]]
* 16:37 James_F: Zuul: Drop REL1_44 testing, EOL, for [[phab:T428911|T428911]]
=== 2026-07-31 ===
* 18:14 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/114
* 12:23 hashar: integration: sudo cumin --force -p 0 'name:docker' 'rm -fR /srv/jenkins/workspace/*pipeline*'
=== 2026-07-30 ===
* 17:28 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover/mediawiki-libs-node-cssjanus/ # [[phab:T424419|T424419]]
=== 2026-07-29 ===
* 21:33 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/109
* 18:33 dancy: Buildkit v0.32.0 deployed to gitlab-cloud-runners staging and production ([[phab:T433520|T433520]])
=== 2026-07-24 ===
* 17:59 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1316011
=== 2026-07-23 ===
* 15:06 dancy: Deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314124 ([[phab:T295351|T295351]])
=== 2026-07-22 ===
* 21:32 dancy: Moved /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/npm to /srv/castor-debug-[[phab:T20260722|T20260722]]-npm-torn-cacache/ on integration-castor06
* 18:41 dancy: Zuul dependencies upgraded and Zuul restarted.
* 18:25 dancy: Zuul is currently broken due to the Gerrit SSH key update. I'm investigating
* 17:24 brennen: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1314014/1 ([[phab:T432886|T432886]])
* 16:58 dancy: Restarting Gerrit ([[phab:T240266|T240266]]) (again)
* 15:59 dancy: Restarting Gerrit ([[phab:T240266|T240266]])
* 09:54 James_F: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure for {{Gerrit|1313321}}, per Lucas_WMDE.
=== 2026-07-21 ===
* 15:59 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1313220
* 07:08 hashar: integration: cleaned up Jenkins workspace on integration-agent-docker-1084
=== 2026-07-20 ===
* 19:36 dancy: Restarting jenkins on contint1003 to clear out hours-long stuck jobs
=== 2026-07-16 ===
* 23:07 mutante: gerrit1003/2002/2003: rm /srv/gerrit/.ssh/config and revert gerrit:1311577 which fixed [[phab:T398401|T398401]] but caused [[phab:T432413|T432413]] - replication is working again
* 18:51 dancy: Updated buildkit to v0.31.2 gitlab-cloud-runners (staging and production) ([[phab:T432360|T432360]])
* 18:37 dancy: Restarted Jenkins to unstick jobs.
* 18:33 dancy: Jenkins placed in shutdown mode in preparation for a restart
* 17:07 bd808: Hard reboot of deployment-cache-text08.deployment-prep.eqiad1.wikimedia.cloud via Horizon; console shows OOM ([[phab:T432374|T432374]])
* 16:40 brennen: contint1002: chown and chmod on /srv/zuul/git/mediawiki/extensions/CampaignEvents/.git per https://www.mediawiki.org/wiki/Continuous_integration/Zuul#Very_high_queue_of_merger:merge_functions
* 16:34 dancy: Restarted docker on contint2003 ([[phab:T432326|T432326]])
* 16:33 dancy: Restarted docker on contint1003 ([[phab:T432326|T432326]])
* 16:32 dancy: Restarted docker on contint1003
* 12:31 Krinkle: krinkle@contint1003:~$ sudo /usr/sbin/service jenkins restart
* 12:10 Krinkle: krinkle@contint1002: zuul restart
=== 2026-07-15 ===
* 18:58 James_F: dev-images: Re-build PHP images for latest point releases, for [[phab:T431099|T431099]]
* 09:13 James_F: Docker: [php83] Update PHP to 8.3.32 for [[phab:T431099|T431099]]
=== 2026-07-14 ===
* 11:53 James_F: Docker: [commit-message-validator] Update to v3.0.0, for [[phab:T431799|T431799]]
* 08:37 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud
=== 2026-07-13 ===
* 15:03 dancy: Updating gitlab-cloud-runners to v19.0.2
=== 2026-07-10 ===
* 11:56 James_F: Zuul: Disable all browser tests on release branches except Wikibase's, for [[phab:T430415|T430415]]
* 09:16 James_F: Docker: [commit-message-validator] Update to 2.3.0
=== 2026-07-09 ===
* 11:55 hashar: retriggering postmerge change for [[phab:T431582|T431582]]: zuul enqueue --trigger gerrit --pipeline postmerge --project machinelearning/liftwing/inference-services --change {{Gerrit|1308631}},3
=== 2026-07-08 ===
* 20:58 brennen: patchdemo: deployed https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/367 ([[phab:T427964|T427964]])
* 19:21 mutante: gerrit - replacing registerEmailPrivateKey in Gerrit config - this invalidates pending/outstanding email validation links for gerrit users - but does not affect active accounts or already verified email addresses
* 14:36 hashar: contint1003, contint2003: manually installed `docker-buildx` Debian package to validate https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308659 # [[phab:T431582|T431582]]
* 13:04 hashar: deployment-prep: git repack on /srv/mediawiki-staging/php-master
* 12:27 hashar: deployment-prep: on deployment server: clearing old branches for mediawiki/extensions and mediawiki/skins # [[phab:T428864|T428864]]
=== 2026-07-07 ===
* 19:41 mutante: contint1003/2003 - add jenkins-agent user to docker group; restart jenkins
* 17:17 hashar: gerrit: deleted /srv/gerrit/java_pid3571660.hprof
=== 2026-07-06 ===
* 20:06 dancy: Updated buildkitd to v0.31.1 in gitlab-cloud-runners ([[phab:T429988|T429988]])
* 08:11 James_F: Zuul: Add WikimediaAntiAbuse extension, for [[phab:T431023|T431023]]
=== 2026-07-03 ===
* 15:25 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add Elastica dep too
* 15:09 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CirrusSearch dep
* 11:22 James_F: Zuul: Make the in-mediawiki-tarball template real
* 10:06 James_F: Zuul: [mediawiki/extensions/TestKitchen] Don't drop from release branches
* 09:53 James_F: Docker: [quibble-coverage] Update phpunit-patch-coverage to 0.0.18, for [[phab:T423987|T423987]] and [[phab:T425807|T425807]]
=== 2026-07-02 ===
* 13:27 James_F: Docker: [composer-scratch] Upgrade composer to 2.10.2 and cascade, for [[phab:T428570|T428570]]
* 10:47 hashar: zuul1002: running Puppet agent to drop `wikimediacloud.org` from `no_proxy` {{!}} [[phab:T430479|T430479]]
=== 2026-07-01 ===
* 13:39 hashar: integration: added timestamping to operations-puppet-catalog-compiler and operations-puppet-catalog-compiler-puppet7-test jobs
* 02:45 hashar: gerrit: on gerrit2003 deleted /srv/gerrit/java_pid3520115.hprof (the JVM apparently died at some point, I assume due to heavy crawling)
=== 2026-06-30 ===
* 15:13 dancy: Rebooting deployment-mwlog02.deployment-prep to clear stuck udp2log processes
=== 2026-06-29 ===
* 09:32 hashar: gerrit: deleted repository phabricator/extensions/BurnDownCharts , created in July 2014, had no commit/changes
=== 2026-06-28 ===
* 15:30 hashar: Updated integration/zuul-jobs from upstream (c75fe6ef19c..fc4af6d4471), notably to remove `requestsexceptions` in `upload-logs-swift` role # [[phab:T430458|T430458]]
=== 2026-06-26 ===
* 13:59 Krinkle: [[phab:T429658|T429658]] krinkle@doc1004:/srv/doc/cover-extensions$ sudo -u doc-uploader rm -rf ShortUrl/
* 13:59 Krinkle: [[phab:T429658|T429658]] krinkle@doc2003:/srv/doc/cover-extensions$ sudo -u doc-uploader rm -rf ShortUrl/
=== 2026-06-25 ===
* 20:57 dancy: Restarting Jenkins to unstick builds
* 20:49 dancy: Investigating castor-save-workspace-cache clog
* 16:50 inflatador: add 60GB cinder vol to deployment-cirrussearch15 [[phab:T425585|T425585]]
* 14:25 inflatador: delete unused servers deployment-cirrussearch1[2-4] [[phab:T425585|T425585]]
* 10:07 James_F: Docker: Bump Node 24 / Node 26 to new releases
=== 2026-06-24 ===
* 17:26 dancy: Set `profile::puppetserver::autosign: /usr/local/sbin/validatecloudvpsfqdn.py` in hiera config for deployment-puppetserver prefix ([[phab:T429413|T429413]])
* 15:12 dancy: sudo systemctl restart php8.3-fpm on deployment-jobrunner05 (Attempting to resolve logspam)
=== 2026-06-23 ===
* 14:31 brennen: deploying patchdemo for https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/361
=== 2026-06-22 ===
* 23:19 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/mediawiki-core/master/mediawiki-node24/ #[[phab:T429824|T429824]]
* 23:07 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ #[[phab:T429824|T429824]] (again)
* 22:35 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ #[[phab:T429824|T429824]]
* 19:30 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-with-gated-extensions-vendor-mysql-php83 #[[phab:T429824|T429824]]
=== 2026-06-19 ===
* 20:13 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1304632
* 09:39 Lucas_WMDE: deployment-deploy04: used createAndPromote to restore User:Lucas Werkmeister (WMDE) to bureaucrat (dewiki, enwiki, metawiki, wikidatawiki) and wikidata-staff (wikidatawiki) after groups were very unhelpfully removed, presumably due to 2FA enforcement, without apparent warning or announcement of any kind
* 07:07 dcausse: restarted php-fpm on deployment-jobrunner05 (inconsistent php state: MediaWiki\Extension\PageAssessments\HookHandler\ParserHooks::__construct(): Argument #3 ($config) must be of type MediaWiki\Config\Config, MediaWiki\Page\WikiPageFactory given)
* 06:13 hashar: Upgrading Quibble jobs to 1.18.2 (retry git connections on reset) https://gerrit.wikimedia.org/r/c/integration/config/+/1304181 # [[phab:T420865|T420865]]
=== 2026-06-18 ===
* 19:52 dancy: Building quibble 1.18.2 images on contint primary ([[phab:T420865|T420865]])
* 18:50 dancy: Tag Quibble 1.18.2 @ {{Gerrit|152b497d317eaafc3cec334d5ce7e549697a2980}} # [[phab:T420865|T420865]]
* 15:15 dancy: Deleting deployment-db11 and deployment-db14 ([[phab:T428910|T428910]])
* 07:23 dcausse: reindexing all wikis to opensearch2 ([[phab:T425585|T425585]], [[phab:T427196|T427196]])
=== 2026-06-17 ===
* 23:16 dduvall: restarted zuul to clear up 5 hrs of stuck queues
* 23:03 mutante: re-enabled puppet on contint1003 - triple checked puppet does NOT start jenkins anymore. BOTH masked AND stopped while running untouched on contint1002. after: gerrit:1303578 {{!}} ([[phab:T418521|T418521]]) ([[phab:T428791|T428791]])
* 18:51 dduvall: performed stop/start of jenkins service on contint1002 following failed safeRestart
* 18:45 dduvall: restarting jenkins due to stuck zuul queues
* 14:56 inflatador: mwscript /srv/mediawiki-staging/php-master/extensions/CirrusSearch/maintenance/ForceSearchIndex.php --wiki=enwikibooks
=== 2026-06-16 ===
* 16:00 dancy: sudo keyholder arm on deployment-deploy04.deployment-prep
* 15:48 dancy: Resizing deployment-deploy04.deployment-prep from g4.cores4.ram8.disk20 to g4.cores8.ram16.disk20 ([[phab:T429364|T429364]])
* 15:31 dancy: Turning off deployment-db11 and deployment-db14
=== 2026-06-15 ===
* 23:40 dancy: systemctl restart php8.3-fpm on deployment-jobrunner05 to reload beta db configuration ([[phab:T428930|T428930]])
* 21:00 dancy: deployment-db15 promoted to master ([[phab:T428930|T428930]])
* 20:53 dancy: deployment-db11.deployment-prep going read-only
* 20:48 dancy: deployment-db11.deployment-prep will be going read-only soon while db15 is being promoted to primary.
* 20:26 dancy: Added deployment-db16 ([[phab:T429245|T429245]])
* 18:26 dancy: Rebooting deployment-jobrunner05
=== 2026-06-12 ===
* 21:45 dancy: Unstuck wmf-beta-update-all service on deployment-deploy04.deployment-prep (sudo systemctl stop wmf-beta-update-all)
* 18:00 thcipriani: unmasking jenkins on contint1002 and restarting
* 17:49 thcipriani: attempting to cancel castor-save-workspace-cache {{Gerrit|6710545}}
* 15:19 James_F: Docker: [php83] Re-platform to Debian Bookworm, for [[phab:T383337|T383337]]
* 15:07 dancy: deployment-db15 configured as a replica of deployment-db11 ([[phab:T428930|T428930]])
* 10:21 Krinkle: `krinkle@<nowiki>{</nowiki>doc1004,doc2003<nowiki>}</nowiki>:/srv/doc/mediawiki-core$ sudo -u doc-uploader rm -rf list/` - remove doc build for git-tag test.
=== 2026-06-11 ===
* 14:46 hashar: for minor in $(seq 21 42); do ./.tox/make-release/bin/python -u ./make-release/branch.py --delete --abandon --bundle '*' "REL1_$minor"; done;
* 14:46 hashar: On all MediaWiki repos, converting old release branches up to REL1_42 included to tags. Last time I missed non wmf repo # [[phab:T380841|T380841]] {{!}} [[phab:T428864|T428864]]
* 09:45 hashar: Converted mediawiki/core branches REL1_39, REL1_40, REL1_41, REL1_42 to tags # [[phab:T428864|T428864]]
* 09:18 hashar: Converting REL1_42 branches to tags # [[phab:T428864|T428864]]
* 09:18 hashar: Converting REL1_41 branches to tags # [[phab:T428864|T428864]]
* 09:05 hashar: Converting REL1_40 branches to tags # [[phab:T428864|T428864]]
* 08:56 hashar: Converting REL1_39 branches to tags # [[phab:T428864|T428864]]
* 08:40 hashar: gerrit: deleted mediawiki/core branch "development" that pointed to {{Gerrit|7f622781cc31053b121f6f4ddbff506cba10d38e}} (which is contained by master). Had probably been created by a direct push.
=== 2026-06-10 ===
* 16:06 James_F: Docker: Provide quibble-bookworm, for [[phab:T362705|T362705]]
=== 2026-06-04 ===
* 23:39 jeena: Updating development images on contint primary for [[phab:T424691|T424691]]
* 12:30 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc # fix failure seen in mwext-node24-rundoc 4812
* 08:35 hashar: Built Docker images `docker-registry.wikimedia.org/releng/java21:0.1` and `docker-registry.wikimedia.org/releng/maven-java21:0.1` # [[phab:T412978|T412978]]
=== 2026-06-03 ===
* 14:19 jnuche: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/107
* 12:54 James_F: Zuul: Add Rae 5e as a trusted user
* 08:26 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1296559 "inference-services: Add LLM generated editing suggestions CI/CD pipelines." # [[phab:T427794|T427794]]
=== 2026-06-02 ===
* 23:23 thcipriani: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/1296695
* 23:10 thcipriani: tag quibble 1.18.1 @ {{Gerrit|4b7959553c095d811426f394165c69ecc13a44eb}}
* 18:13 brennen: devtools phab/phorge: deployed work/2026-06-01-merge-phorge to https://phabricator.wmcloud.org/ for testing ([[phab:T410849|T410849]])
* 16:05 hashar: integration: upgraded pypy from 7.3.11 (py3.9) to 7.3.20 (py3.11) # [[phab:T423607|T423607]]
* 15:49 jnuche: Updating buildkitd to v0.30.0 in gitlab-cloud-runners ([[phab:T426212|T426212]])
* 15:24 hashar: Building docker images for https://gerrit.wikimedia.org/r/c/integration/config/+/1295865 # [[phab:T423607|T423607]]
* 14:57 jnuche: Jenkins/Zuul is back
* 14:38 jnuche: restarting Jenkins
* 14:28 jnuche: bring back castor node, that didn't help
* 14:23 jnuche: trying to reconnect castor node, see if that helps somehow
* 14:12 jnuche: option to "Enable Gearman" times out. Can't re-enable from UI. Gearman plugin logs are empty. Neat
* 14:05 jnuche: trying to reconnect Gearman
=== 2026-06-01 ===
* 21:34 jeena: Updating development images on contint primary for [[phab:T424663|T424663]]
* 10:58 hashar: gerrit: flushed `ldap_usernames` cache in case a missing account ended up being cached there # [[phab:T427792|T427792]]
=== 2026-05-28 ===
* 16:48 hashar: castor: nuked SonarQube cache: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-codehealth-master-non-voting/sonar/ # [[phab:T427471|T427471]]
* 07:05 hashar: integration: delete deployment-deploy04 agent from Jenkins controller. Batch job has been migrated to a systemd driven script # [[phab:T256168|T256168]]
* 06:29 hashar: integration: delete integration-castor05 agent from Jenkins, replaced by integration-castor06 # [[phab:T421114|T421114]]
=== 2026-05-27 ===
* 21:07 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1294425
* 09:49 hashar: Updated helm-lint Jenkins job to use releng/helm-linter:0.8.0 image # [[phab:T424824|T424824]]
* 08:25 codders: integration-castor06: sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/
=== 2026-05-26 ===
* 15:26 James_F: Zuul: Add PHP 8.5 to a few missing PHP pipelines, oops
* 14:46 hashar: Updating Quibble Jenkins jobs to stop Supervisord from spawning Memcached # [[phab:T397810|T397810]]
* 14:07 hashar: Updated integration-quibble-* jobs in order to validate running tests with Supervisord not managing memcached # [[phab:T397810|T397810]]
* 13:15 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-docs-publish # fix failure seen in mwext-node24-docs-publish 981, 985, 987
* 09:05 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1292620 (Introduce Phan composer job - [[phab:T231966|T231966]])
=== 2026-05-21 ===
* 16:33 hashar: Reloaded Zuul to enable Node24 CI job for `labs/tools/wdaudiolex-fe` # [[phab:T426366|T426366]]
=== 2026-05-19 ===
* 18:07 James_F: Zuul: [mediawiki/libs/ZestJQ] Add basic PHP and Node CI
=== 2026-05-18 ===
* 21:12 James_F: Zuul: [operations/software/gerrit] Add Node26 as experimental
* 21:10 mutante: gerrit-replica.wikimedia.org, gerrit-spare.wikimedia.org - rebooting backends
* 20:57 James_F: Zuul: [integration/docroot] Test in PHP 8.3+, dropping 8.2
* 20:56 James_F: Zuul: [analytics/wmde/scripts] Test in PHP 8.3+, dropping 8.
* 20:18 James_F: Zuul: [wikimedia/fundraising/dash] Replace Node 20 testing with Node 24
* 20:18 James_F: Zuul: Migrate various labs things to Node 24
* 20:02 James_F: Docker: [ajv, sonar-scanner] Migrate to Node 24
* 19:58 James_F: Zuul: Migrate various production/CI things to Node 24
* 18:17 mutante: releases.wikimedia.org - rebooting backends
* 18:14 mutante: rebooting production gitlab-runners
* 18:12 dancy: gitlab-cloud-runners have been revived.
* 18:11 James_F: Zuul: [design/codex] Switch CI to Node 24
* 15:52 dancy: gitlab-cloud-runners are in a broken state. I'm investigating
* 14:27 hashar: Upgrading Quibble jobs to 1.18.0
* 09:29 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure seen in mediawiki-node24 22260
=== 2026-05-15 ===
* 18:36 dancy: Upgraded gitlab-cloud-runners (prod) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 18:24 dancy: Upgrading gitlab-cloud-runners (prod) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 18:11 dancy: Upgraded gitlab-cloud-runners (staging) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 17:59 dancy: Upgrading gitlab-cloud-runners (staging) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 13:02 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1287828 [[phab:T426392|T426392]]
=== 2026-05-13 ===
* 12:42 James_F: Zuul: [mediawiki/extensions/Springboard] Add AdminLinks Phan dependency
* 12:42 James_F: Zuul: [mediawiki/extensions/ChatBot] Add dependencies on VisualEditor and BlueSpiceFoundation
* 12:42 James_F: Zuul: [mediawiki/extensions/ChatIntegration] Add dependency on VisualEditor
* 12:37 James_F: Zuul: [mediawiki/extensions/WikiLambda] Drop AF and SB deps down to phan-only, for [[phab:T423180|T423180]]
=== 2026-05-12 ===
* 20:57 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/104 ([[phab:T424774|T424774]])
* 18:08 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add AF and SB deps for [[phab:T423180|T423180]]
* 14:18 atsukoito: PrivateSettings: empty $wgOpensearchCredentials for opensearch-on-k8s synced to deploy04 by Reedy
* 13:04 atsukoito: PrivateSettings: credentials for opensearch-on-k8s ttmserver-test
* 11:50 James_F: Zuul: [machinelearning/liftwing/inference-services] Add qwen36 llm model CI/CD pipelines, for [[phab:T425680|T425680]]
* 11:46 James_F: Zuul: Add experimental php-pie-build* jobs to other PHP extensions, for [[phab:T425943|T425943]]
* 11:37 James_F: Zuul: [mediawiki/php/wikidiff2] Add experimental php-pie-build* jobs, for [[phab:T425943|T425943]]
* 10:05 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-with-Wikibase-extensions-browser-tests-only-vendor-php83 # fix failure seen in quibble-with-Wikibase-extensions-browser-tests-only-vendor-php83 7817
* 08:44 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # broken Cypress cache? hopefully fix failure seen in quibble-vendor-mysql-php83-selenium 51633
=== 2026-05-11 ===
* 18:28 James_F: Docker: Add changes to php-compile images for PIE, for [[phab:T425943|T425943]]
* 16:06 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # broken Cypress cache? hopefully fix failure seen in quibble-vendor-mysql-php83-selenium 51439 and 51452
=== 2026-05-09 ===
* 20:46 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1285498
=== 2026-05-07 ===
* 22:53 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/105
=== 2026-05-06 ===
* 18:13 bd808: Unblock 88.165.192.0/19
* 18:03 bd808: Unblock 94.208.0.0/14
* 17:56 bd808: Unblock 84.226.0.0/16
* 17:41 bd808: Unblock 94.34.0.0/16
* 17:35 bd808: Unblock 109.134.0.0/16
=== 2026-05-05 ===
* 21:20 James_F: Zuul: Provide Node 26 experimental jobs everywhere needed
* 21:04 James_F: Docker: Provide initial Node 26 images
* 19:01 James_F: Zuul: [mediawiki/extensions/PageAssessments] Add Scribunto dependency, for [[phab:T396135|T396135]]
* 14:58 dancy: rm /var/log/<nowiki>{</nowiki>user.log.1,syslog.1,messages.1<nowiki>}</nowiki> on deployment-eventgate-4.deployment- prep ([[phab:T425429|T425429]])
=== 2026-05-04 ===
* 15:19 dancy: Upgrading gitlab cloud runners (prod) from 1.35.1-do.3 to 1.35.1-do.5
* 14:51 dancy: Upgrading gitlab cloud runners (staging) from 1.35.1-do.3 to 1.35.1-do.5
* 10:40 James_F: Zuul: Provide non-voting PHP 8.4/8.5 Quibble jobs for bluespice template
=== 2026-05-02 ===
* 20:49 James_F: Zuul: [mediawiki/core] Enforce PHP 8.4 & 8.5 on release branches, all pass
* 19:27 James_F: Zuul: Provide non-voting PHP 8.4/8.5 Quibble jobs for MW release branches
* 19:19 James_F: Zuul: [mediawiki/extensions/BlogPage] Add dependencies
* 16:48 James_F: Hard-restarting Zuul to clear the huge number of i18n updates being re-submitted.
* 15:48 James_F: Zuul: [wikimedia-cz/*] Test in PHP 8.3+, dropping 8.2
* 14:02 TheresNoTime: Add bvibber to deployment-prep project
* 09:08 James_F: Docker: [quibble-*] Add php-luasandbox so we can test both modes in Scribunto
=== 2026-05-01 ===
* 15:42 James_F: Zuul: [wikimedia/lucene-explain-parser] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: Zuul: [wikimedia/textcat] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: Zuul: [mediawiki/tools/ParseWiki] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: zuul: Add ToprakM to CI allowlist
* 15:19 James_F: Zuul: [translatewiki] Test in PHP 8.3+, dropping 8.2
* 15:10 James_F: Zuul: [mediawiki/extensions/WikiEditor] Add TestKitchen as a dependency, for [[phab:T425076|T425076]]
* 12:40 James_F: Zuul: [mediawiki/tools/code-utils] Test in PHP 8.3+, dropping 8.2
* 08:02 James_F: Zuul: Update xtex's e-mail in the allowlist
* 07:37 James_F: Zuul: Switch release branches' selenium jobs to PHP 8.3
* 07:33 James_F: Zuul: Test Wikimedia production libraries in PHP 8.3+, dropping 8.2
=== 2026-04-30 ===
* 21:36 brennen: gitlab-webhooks: building & restarting to deploy https://gitlab.wikimedia.org/repos/releng/gitlab-webhooks/-/merge_requests/40
* 20:26 James_F: Zuul: [mediawiki/tools/api-testing] Make PHP 8.5 CI voting
* 20:16 James_F: jforrester@doc1004:~$ # sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/WebAuthn/ # [[phab:T415832|T415832]]
* 20:14 James_F: Zuul: [mediawiki/extensions/WebAuthn] Archive, for [[phab:T415832|T415832]] / [[phab:T303495|T303495]]
* 17:16 brennen: wikibugs: most maintainers at hackathon, so go release-engineering added as a maintainer while looking to debug error at https://gitlab.wikimedia.org/toolforge-repos/wikibugs2/-/jobs/810904
* 15:19 mutante: upgrading zuul to 14.2.0-1 on "new zuul" machines ([[phab:T424879|T424879]])
=== 2026-04-29 ===
* 15:49 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add ConfirmEdit dependency, for [[phab:T424597|T424597]]
* 15:36 James_F: Zuul: Drop experimental node22 jobs, never used in practice
* 15:28 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1279392, https://gerrit.wikimedia.org/r/1279397
=== 2026-04-28 ===
* 18:11 bd808: Unblock 86.0.0.0/16
* 17:41 bd808: Unblock 79.192.0.0/10
* 17:07 James_F: Zuul: [mediawiki/tools/phpunit-patch-coverage] Drop PHP 8.2 testing
* 17:07 James_F: Zuul: [mediawiki/tools/minus-x] Drop PHP 8.2 testing
* 17:07 James_F: Zuul: [mediawiki/tools/codesniffer] Drop PHP 8.2 testing
* 16:32 James_F: Zuul: [mediawiki/services/jobrunner] Drop PHP 8.2 testing
* 13:34 James_F: Zuul: [mediawiki/tools/phan] Drop PHP 8.2 testing
* 13:34 James_F: Zuul: [oojs/ui] Drop PHP 8.2 testing
* 13:14 James_F: Zuul: [mediawiki/tools/phan/SecurityCheckPlugin] Drop PHP 8.2 CI
* 10:40 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud to fix failure seen in mwext-node24-rundoc #1717
* 00:03 bd808: Increase parallelism for wmf-beta-update-databases.py ([[phab:T256168|T256168]])
=== 2026-04-27 ===
* 22:11 bd808: Beta Cluster MediaWiki update logs now available via https://beta-update.wmcloud.org/ ([[phab:T256168|T256168]])
* 21:57 bd808: Add web security group to deployment-deploy04 ([[phab:T256168|T256168]])
* 20:45 James_F: Zuul: Restrict mw*-codehealth-patch jobs to master only, for [[phab:T424573|T424573]]
* 17:16 James_F: Docker: [mediawiki-phan-taint-check-demo] Re-platform to Trixie and so PHP 8.4
* 15:53 James_F: Zuul: [mediawiki/extensions/ReportIncident] Add TestKitchen phan dependency, for [[phab:T424220|T424220]]
* 14:32 James_F: Zuul: Drop PHP 8.2 enforcement from MediaWiki things for master and REL1_46 for [[phab:T358667|T358667]]
* 12:38 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-docs-publish # fix failure seen in mwext-node24-docs-publish 383
* 09:18 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover/mediawiki-libs-node-cssjanus/ # [[phab:T424419|T424419]]
=== 2026-04-26 ===
* 20:49 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1276777
=== 2026-04-24 ===
* 22:48 dduvall: merged zuul3 branch of integration/config into master and pushed (in preparation for https://gerrit.wikimedia.org/r/c/operations/puppet/+/1277198)
* 12:27 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1276428
=== 2026-04-23 ===
* 23:57 bd808: Set `profile::beta::autoupdater::run_updater: true` for deployment-deploy04 via Horizon ([[phab:T256168|T256168]])
* 22:58 bd808: bd808@deployment-deploy04 `sudo -u jenkins-deploy /usr/local/bin/wmf-beta-update-all`
* 22:36 bd808: bd808@deployment-deploy04 `sudo -u mwdeploy /usr/local/bin/wmf-beta-update-all`
* 22:16 bd808: Disabled https://integration.wikimedia.org/ci/view/Beta/job/beta-update-databases-eqiad so that replacement script can be tested ([[phab:T256168|T256168]])
* 22:12 bd808: Disabled https://integration.wikimedia.org/ci/job/beta-code-update-eqiad so that replacement script can be tested ([[phab:T256168|T256168]])
* 22:02 bd808: Cherry-picked {{gerrit|1276813}} to deployment-puppetserver-1 ([[phab:T256168|T256168]])
* 20:11 James_F: Zuul: [wikibase/*] Replace CI testing in Node 20 with Node 24
* 20:11 James_F: Zuul: [wikidata/query/*] Replace CI testing in Node 20 with Node 24
* 20:11 James_F: Zuul: [analytics/*] Replace CI testing in Node 20 with Node 24
* 20:10 James_F: Zuul: [mediawiki/tools/*] Replace CI testing in Node 20 with Node 24
* 20:06 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.34.5-do.3 to 1.35.1-do.3 ([[phab:T423726|T423726]])
* 19:55 James_F: Zuul: [jquery-client] Replace CI testing in Node 20 with Node 24
* 19:51 James_F: Zuul: [wikipeg] Drop testing in Node 20 and Node 22
* 19:47 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.34.5-do.3 to 1.35.1-do.3 ([[phab:T423726|T423726]])
* 19:37 James_F: Zuul: [oojs/ui] Drop CI testing in Node 20 and Node 22
* 19:37 James_F: Zuul: [oojs/js] Drop CI testing in Node 20 and Node 22
* 19:37 James_F: Zuul: [unicodejs] Replace CI testing in Node 20 with Node 24
* 19:36 James_F: Zuul: [wikimedia/portals] Drop CI testing in Node 20 and Node 22
* 18:57 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.33.9-do.3 to 1.34.5-do.3 ([[phab:T423726|T423726]])
* 18:39 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.33.9-do.3 to 1.34.5-do.3 ([[phab:T423726|T423726]])
* 18:18 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.33.9-do.2 to 1.33.9-do.3 ([[phab:T423726|T423726]])
* 17:58 James_F: Zuul: [mediawiki/extensions/OAuth] Add dependency on CentralAuth, for [[phab:T415281|T415281]]
* 17:56 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.32.13-do.2 to 1.33.9-do.3 ([[phab:T423726|T423726]])
* 16:35 James_F: Zuul: Enforce PHP 8.5 CI for MW things in master (and REL1_46), for [[phab:T411814|T411814]]
* 16:19 James_F: Zuul: [mediawiki/services/parsoid] Enable PHP 8.5 CI
* 15:47 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add AntiSpoof dependency, for [[phab:T420548|T420548]]
* 14:20 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure seen in mediawiki-node24 8385 and 8405
* 12:56 James_F: Zuul: [mediawiki/extensions/GrowthExperiments] Add CentralNotice dependency, for [[phab:T422082|T422082]]
=== 2026-04-22 ===
* 00:07 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add MF dependency, for [[phab:T424113|T424113]]
=== 2026-04-21 ===
* 23:26 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CommunityConfiguration dep too, for [[phab:T394410|T394410]]
* 23:17 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add standalone test jobs, for [[phab:T422031|T422031]]
* 20:47 inflatador: updating cirrussearch hosts to Trixie/OpenSearch 2 [[phab:T421763|T421763]]
* 20:38 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CommunityConfiguration phan dep, for [[phab:T394410|T394410]]
* 20:17 bd808: Running tofu for [[phab:T421244|T421244]]
* 18:00 James_F: Zuul: [mediawiki/extensions/WatchAnalytics] Add ApprovedRevs Phan dependency
* 16:35 bd808: Unblock 79.116.0.0/16
* 13:34 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add TestKitchen phan dep, for [[phab:T415254|T415254]]
* 13:27 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add CentralAuth dependency, for [[phab:T420548|T420548]]
=== 2026-04-20 ===
* 23:56 bd808: Unblock 76.157.0.0/16
* 18:28 dancy: Upgrading gitlab cloud runners (staging) to 1.33.9-do.2 ([[phab:T423726|T423726]])
* 18:28 dancy: Upgrading gitlab cloud runners (staging) ([[phab:T423726|T423726]])
* 18:19 James_F: jjb: All 486 (!) jobs now updated for [[phab:T423622|T423622]]
* 18:18 bd808: Unblock 113.128.0.0/15
* 15:03 James_F: Docker: Bump ci-bullseye/-bookworm/-trixie for mirrors.wm.org removal, [[phab:T423622|T423622]]
=== 2026-04-19 ===
* 19:53 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1272752
=== 2026-04-17 ===
* 21:07 thcipriani: marking integration-agent-1080 offline for experimentation
* 19:30 thcipriani: reconfiguring castor-save-workspace-cache with https://gerrit.wikimedia.org/r/1273935
* 17:47 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.32.10-do.1 to 1.32.13-do.2 ([[phab:T423726|T423726]])
* 16:49 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.32.10-do.1 to 1.32.13-do.2 ([[phab:T423726|T423726]])
=== 2026-04-16 ===
* 20:49 dduvall: creating integration/zuul-jobs repo to serve as a mirror of opendev.org/zuul/zuul-jobs ([[phab:T406384|T406384]])
* 13:38 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1272711 [[phab:T423568|T423568]]
* 11:07 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud
=== 2026-04-15 ===
* 20:05 James_F: Zuul: Configure REL1_46 CI, for [[phab:T423257|T423257]]
* 17:44 bd808: Unblock 176.0.0.0/13
* 17:39 bd808: Unblock 46.128.0.0/16
* 17:32 bd808: Unblock 176.86.0.0/16
* 16:39 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/commit/127d783b2176ac60b646a5fa4f1b1a872ca66340
* 15:33 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/100
* 01:02 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/99
=== 2026-04-14 ===
* 20:42 James_F: Docker: [composer-scratch] Upgrade composer to 2.9.7 and cascade
* 16:35 bd808: Unblock 88.112.0.0/14
* 00:48 bd808: Unblock 24.6.0.0/16
* 00:42 bd808: Unblock 152.231.48.0/20
=== 2026-04-13 ===
* 22:00 James_F: Zuul: [mediawiki/vendor] Drop accidental Wikibase browser tests on branches
* 20:28 James_F: Zuul: [mediawiki/extensions/Chart] Drop Doxygen publish job, not used
* 14:42 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add FlaggedRevs dep, for [[phab:T421011|T421011]]
=== 2026-04-12 ===
* 18:21 James_F: jforrester@contint1002:~$ sudo /usr/sbin/service zuul restart && tail -f -n100 /var/log/zuul/zuul.log # [[phab:T423027|T423027]]
=== 2026-04-10 ===
* 23:22 James_F: jforrester@contint1002:~$ zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/extensions/ReadingLists --change {{Gerrit|1269498}},2 # [[phab:T422976|T422976]]
* 23:20 James_F: Zuul: [mediawiki/extensions/ReadingLists] Publish JS coverage, for [[phab:T422976|T422976]]
* 23:13 James_F: Zuul: Migrate a few straggler Node 20 MediaWiki things to Node 24
* 23:01 James_F: Zuul: Move all MediaWiki things from mediawiki-node20 to mediawiki-node24
* 21:59 James_F: Docker: Bump Node base images to March releases and cascade; Upgrade Quibble images from Node 20 to Node 24
* 10:24 hashar: Updating all Quibble jobs to 1.17.1
* 10:22 hashar: Updated PostgreSQL jobs to Quibble 1.17.1 # [[phab:T422110|T422110]]
* 10:22 hashar: Updated apitesting job to Quibble 1.17.1 # [[phab:T422843|T422843]] [[phab:T418743|T418743]]
* 09:51 hashar: Tag Quibble 1.17.1 @ {{Gerrit|0a1ab3b7c3dfee36c9bc2e9b049957d94e190e85}}
=== 2026-04-09 ===
* 15:13 hashar: Rolling back Quibble jobs to 1.16.0 (api-testing stage fails due to missing npm install step`
* 14:58 hashar: Upgrading Quibble jobs to 1.17.0
* 14:23 hashar: Tagged Quibble 1.17.0 @ {{Gerrit|864381c6b63bdbcd8c74a3162c406fffcaaf8694}}
* 07:48 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1268559 "Zuul: use standalone jobs for GrowthExperiments Cypress tests" {{!}} [[phab:T417412|T417412]]
=== 2026-04-08 ===
* 22:19 dancy: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1269068
* 22:01 bd808: Unblock 95.216.12.170/32 ([[phab:T422751|T422751]])
* 19:26 brennen: gitlab-webhooks: building & deploying https://gitlab.wikimedia.org/repos/releng/gitlab-webhooks/-/merge_requests/37 - hitting some build tooling stuff, trying a fix per instructions in the error log
* 17:54 bd808: Unblock 167.56.0.0/13 ([[phab:T422721|T422721]])
* 06:31 hashar: Deleted integration-agent-castor05 Bullseye instance, replaced by integration-agent-castor06 which is on Bookworm # [[phab:T421114|T421114]]
* 06:24 hashar: Deleted integration-agent-qemu-1003 Bullseye image, replaced by integration-agent-qemu-1004 which is on Bookworm # [[phab:T422488|T422488]]
=== 2026-04-07 ===
* 22:25 dduvall: adding new pipelinelib labels to ci nodes ([[phab:T422234|T422234]])
* 20:05 hashar: Triggered a build of https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/
* 17:06 dduvall: added `Docker` label to `contint` jenkins nodes ([[phab:T422507|T422507]])
* 17:05 dduvall: restored missing `pipelinelib` labels on `integration-agent-docker-` CI hosts ([[phab:T422507|T422507]])
* 16:53 bd808: Unblock 73.0.0.0/8 ([[phab:T422498|T422498]])
* 12:36 hashar: jjb: use $CASTOR_HOST for Quibble success cache. https://gerrit.wikimedia.org/r/1268545 {{!}} This causes the Quibble jobs to use a new instance for the success cache, which is empty # [[phab:T383243|T383243]] [[phab:T421114|T421114]]
* 12:17 hashar: Migrated Castor from integration-castor05 to integration-castor06. Updated CASTOR_HOST in Jenkins and moved the Cinder volume to the new instance # [[phab:T421114|T421114]]
* 11:14 hashar: Added Bookworm based Jenkins agents to the pool Hostnames 1090, 1091, 1092 and 1093 # [[phab:T421114|T421114]]
* 10:09 hashar: Added Bookworm based Jenkins agents to the pool Hostnames 1083 to 1089 # [[phab:T421114|T421114]]
* 07:23 hashar: CI Jenkins: removed `blubber` label from all agents after having moved PipelineLib to use the `Docker` label {{!}} [[phab:T422234|T422234]]
=== 2026-04-06 ===
* 16:01 dancy: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1268239
=== 2026-04-03 ===
* 20:17 bd808: Unblock 2.54.0.0/16 ([[phab:T422238|T422238]])
* 17:25 bd808: Unblock 31.18.0.0/16 ([[phab:T422245|T422245]])
* 17:18 bd808: Unblock 2.54.128.0/19 ([[phab:T422238|T422238]])
* 16:18 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1264649 "add Python 3.14 to pywikibot jobs and separate lint tests" {{!}} [[phab:T421723|T421723]]
* 09:26 hashar: integration: nuked pywikibot/core pre-commit cache # [[phab:T422242|T422242]]
* 09:15 hashar: Added Bookworm based Jenkins agents to the pool with label `Docker`. Hostnames are `integration-agent-docker-107*` # [[phab:T421114|T421114]]
* 02:47 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1267398
=== 2026-04-02 ===
* 16:50 thcipriani: restart jenkins
* 15:15 bd808: Unblock 82.216.0.0/16 ([[phab:T421508|T421508]])
* 15:07 bd808: Unblock 95.90.0.0/15 ([[phab:T421485|T421485]])
* 11:19 James_F: Zuul: [oojs/ui] Drop ooui-ruby2.7-rake job, we're abandoning Ruby use there
=== 2026-04-01 ===
* 22:01 bd808: Unblock 109.144.0.0/12 ([[phab:T422019|T422019]])
* 20:16 bd808: Unblock 93.192.0.0/10 ([[phab:T421894|T421894]])
* 19:25 dancy: Updating buildkitd to v0.29.0 in gitlab-cloud-runners (prod) ([[phab:T415284|T415284]])
* 17:57 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/97 ([[phab:T420441|T420441]])
* 17:39 bd808: Unblock 94.134.0.0/15 ([[phab:T421866|T421866]])
* 16:31 dancy: Upgrade buildkit to 0.29.0 in staging gitlab-cloud-runners ([[phab:T415284|T415284]])
* 10:47 taavi: integration-castor05: free up a bit of disk space by deleting cache for AhoCorasick/ CLDRPluralRuleParser/ HtmlFormatter/ RelPath/ RunningStat/ IPSet/
=== 2026-03-30 ===
* 22:01 bd808: Unblock 78.20.0.0/14 ([[phab:T421586|T421586]])
* 21:04 bd808: Unblock 95.88.0.0/15 ([[phab:T421774|T421774]])
* 20:49 bd808: Unblock 95.89.191.0/24 ([[phab:T421774|T421774]])
* 20:29 bd808: Unblock 73.162.0.0/16 ([[phab:T421549|T421549]])
* 13:10 hashar: gerrit: abandon mediawiki/core changes that are 2+years old and are attached to a task (`Bug: Txxxx`)
* 11:37 hashar: Reloaded Zuul to to add 3 persons to the allow list
* 10:43 James_F: Docker: Re-pushing to try to create quibble-coverage 1.16.0-s2
=== 2026-03-27 ===
* 21:00 James_F: Docker: [quibble-bullseye] Drop Python 2 from images
* 11:28 hashar: deployment-prep: removed block for `143.176.0.0/15` and blocked subblock `143.176.0.0/16` instead. This unblocks `143.177.0.0/16` # [[phab:T421420|T421420]]
* 00:18 bd808: Unblock 95.90.238.0/23 ([[phab:T421447|T421447]])
=== 2026-03-26 ===
* 21:25 bd808: Unblock 89.240.0.0/15 ([[phab:T421364|T421364]])
* 21:09 brennen: patchdemo: deploy to production for https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/312
=== 2026-03-25 ===
* 20:41 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256318 [[phab:T421283|T421283]]
* 15:46 dancy: Migrated gitlab-cloud-runners (prod) from nginx-ingress to traefik ([[phab:T420743|T420743]])
* 15:32 dancy: Migrated gitlab-cloud-runners (staging) from nginx-ingress to traefik ([[phab:T420743|T420743]])
* 10:01 hashar: Updating tox Jenkins jobs to add support for Python 3.14 {{!}} https://gerrit.wikimedia.org/r/1260632 {{!}} [[phab:T421209|T421209]]
* 08:40 codders: integration: integration-castor05: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20/
=== 2026-03-24 ===
* 19:40 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1255746
* 15:34 brennen: gitlab1004: manual test run of `configure-projects` with cleared issue allowlist ([[phab:T412882|T412882]])
* 15:26 bd808: Unblock 47.194.0.0/16 ([[phab:T421127|T421127]])
* 12:53 hashar: integration: deleted old Puppet 5 compiler agents from Jenkins ( pcc-worker1014.puppet-diffs.eqiad1.wikimedia.cloud , pcc-worker1015.puppet-diffs.eqiad1.wikimedia.cloud , pcc-worker1016.puppet-diffs.eqiad1.wikimedia.cloud ) # [[phab:T367399|T367399]]
* 07:42 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1259755
=== 2026-03-23 ===
* 15:28 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 90272
=== 2026-03-22 ===
* 14:52 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1258082
* 01:00 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256488
=== 2026-03-21 ===
* 08:10 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256962
* 07:48 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256946
=== 2026-03-20 ===
* 21:21 bd808: Unblock 103.159.218.0/24 ([[phab:T420530|T420530]])
* 14:59 James_F: Zuul: [mediawiki/extensions/AbuseFilter] Add dependency on CodeMirror, for [[phab:T399673|T399673]]
=== 2026-03-19 ===
* 16:54 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1255777
* 16:01 Krinkle: Hoist l10n-bot rights from labs/tools parent to labs parent to reduce duplication in other labs/ repos
* 15:50 Krinkle: Create labs/xtools repo (branch: main, parent: labs, owner: labs-xtools), ref [[phab:T402086|T402086]]
=== 2026-03-18 ===
* 21:11 dcausse: [[phab:T403775|T403775]]: reindexing all wikis to enable new sorting options
* 21:08 dcausse: restarting opensearch on deployment-cirrussearch(12{{!}}13{{!}}14) instances to pickup new plugin versions
* 14:56 James_F: Zuul: Handle wmf/next the same way as wmf/branch_cut_pretest
* 14:52 James_F: Zuul: [GrowthExperiments] drop duplicate VisualEditor dep
* 14:52 James_F: Zuul: [search/*] Add experimental Java 25 jobs
=== 2026-03-17 ===
* 22:50 James_F: Zuul: [mediawiki/extensions/JsonForms] Add quibble jobs
* 21:27 James_F: Zuul: search: Update opensearch plugins for Java 11/17, for [[phab:T420407|T420407]]
* 20:20 bd808: Resize deployment-sessionstore06 from g4.cores1.ram2.disk20 to g4.cores2.ram4.disk20 ([[phab:T415021|T415021]])
* 16:43 James_F: Zuul: [BlueSpicePermissionManager] Add …ConfigManager & …UserManager deps
* 14:36 James_F: Zuul: [mediawiki/extensions/ArticleGuidance]: Add SpamBlacklist as phan dep, for [[phab:T420015|T420015]]
=== 2026-03-13 ===
* 13:59 andrewbogott: deleting ptr record 117.0.16.172.in-addr.arpa. -- accidental duplicate for deployment-kafka-logging01.deployment-prep.eqiad1.wikimedia.cloud
* 13:04 elukey: re-create kafka-logging-01 in deployment-prep on trixie and Kafka 3.7 (was running on buster)
* 09:13 elukey: upgrade kafka-jumbo and kafka-main to Confluent 7.7 in deployment-prep (pre-requisite before being able to upgrade to Trixie)
=== 2026-03-12 ===
* 21:23 bd808: Hard reboot deployment-sessionstore06 ([[phab:T415021|T415021]])
* 01:14 James_F: Docker: [helm-linter] Bump for Envoy 1.35.9, for [[phab:T419637|T419637]]
=== 2026-03-11 ===
* 16:48 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/MetricsPlatform # [[phab:T417568|T417568]]
* 16:47 James_F: Zuul: [mediawiki/extensions/MetricsPlatform] Archive, for [[phab:T416865|T416865]]
* 11:12 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1250529 "inference-services: Split policy violation CI into separate model jobs." - [[phab:T418832|T418832]]
=== 2026-03-10 ===
* 17:39 dduvall: deployed reggie v1.18.0 to gitlab-cloud-runner production
* 17:11 hashar: Updated MediaWiki coverage jobs so that they now keep "Generate a local configuration by running `composer phpunit:config`" message # [[phab:T419073|T419073]]
* 16:41 dduvall: deployed reggie v1.18.0 to gitlab-cloud-runner staging
* 08:21 codders: integration: integration-castor05: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20
=== 2026-03-09 ===
* 21:53 bd808: Reboot deployment-shellbox01 on the off chance that is makes the new permissions error go away ([[phab:T419440|T419440]])
* 13:13 James_F: Zuul: [mediawiki/extensions/WikiShare] Mark as archived, for [[phab:T413589|T413589]]
* 13:11 James_F: Zuul: [mediawiki/extensions/Memento] Mark as archived, for [[phab:T369991|T369991]]
* 13:10 James_F: Zuul: [mediawiki/extensions/QuickGV] Mark as archived, for [[phab:T413348|T413348]]
* 13:10 James_F: Zuul: [mediawiki/extensions/SemanticImageInput] Mark as archived, for [[phab:T413588|T413588]]
* 13:09 James_F: Zuul: [mediawiki/extensions/SidebarDonateBox] Mark as archived, for [[phab:T413587|T413587]]
* 13:07 James_F: Zuul: [mediawiki/extensions/SemanticSifter] Mark as archived, for [[phab:T413586|T413586]]
* 13:06 James_F: Zuul: [mediawiki/extensions/GoogleAdSense] Mark as archived, for [[phab:T413585|T413585]]
* 13:04 James_F: Zuul: [mediawiki/extensions/SecurityAPI] Mark as archived, for [[phab:T418008|T418008]]
* 12:50 James_F: Zuul: [mediawiki/extensions/CheckUser] Add DiscussionTools dependency
* 12:50 James_F: Zuul: [mediawiki/skins/MinervaNeue] Add dependencies for TestKitchen
* 10:40 hashar: gerrit: mediawiki/vendor: converted `es6` and `es710` branches to tags # [[phab:T417804|T417804]]
* 09:24 hashar: Updating Quibble jobs to 1.16.0 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1248880 {{!}} [[phab:T417399|T417399]] [[phab:T417409|T417409]] [[phab:T418461|T418461]]
* 09:15 hashar: updating all CI Jenkins jobs using `./jjb-update`
=== 2026-03-06 ===
* 19:46 James_F: Zuul: [mediawiki/services/geoshapes] Mark as archived, for [[phab:T418372|T418372]]
* 16:37 hashar: Building Docker images for Quibble 1.16.0
* 16:31 hashar: Tag Quibble 1.16.0 @ {{Gerrit|0b9db5fe3cabb2cec0b5d44e128bafa917b3b895}} # [[phab:T417399|T417399]] [[phab:T417409|T417409]] [[phab:T418461|T418461]]
* 12:32 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1248411 "jjb, Zuul: vary Wikibase Selenium for release branches" {{!}} [[phab:T418797|T418797]]
* 12:12 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1248409/ "jjb, Zuul: rename wikibase-selenium job for clarity" {{!}} [[phab:T418797|T418797]]
=== 2026-03-05 ===
* 14:41 James_F: Zuul: [mediawiki/skins/MinervaNeue] Add TestKitchen as a dependency for [[phab:T418053|T418053]]
* 08:01 hashar: Reloaded Zuul to rename wikibase-client / wikibase-repo jobs {{!}} https://gerrit.wikimedia.org/r/1238317
* 00:04 James_F: Docker: [quibble-coverage] Use local PHPUnit config, for [[phab:T345481|T345481]]
=== 2026-03-04 ===
* 21:16 James_F: Zuul: [mediawiki/core] Make PHP 8.5 voting on master branch, for [[phab:T411814|T411814]]
* 21:10 James_F: Zuul: [mediawiki/vendor] Make PHP 8.5 voting on master branch, for [[phab:T411814|T411814]]
* 19:48 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/96 ([[phab:T419004|T419004]])
* 18:50 James_F: Revert "Zuul: [mediawiki/extensions/MobileFrontend] Add ParserMigration dependency", for [[phab:T419043|T419043]]
* 16:23 James_F: Zuul: [mediawiki/services/parsoid] Make PHP 8.4 voting
* 15:37 James_F: Docker: [rake-ruby2.7] Add libffi-dev too, for [[phab:T418463|T418463]]
* 13:59 James_F: Docker: [rake-ruby2.7] Add ruby-ffi for [[phab:T418463|T418463]]
* 13:54 hashar: SIGKILL Zuul cause it can't gracefully stop most probably due to being locked attempting to report back to Gerrit # [[phab:T419009|T419009]]
* 13:49 hashar: Stopping Zuul # [[phab:T419009|T419009]]
* 13:41 hashar: Took a Zuul stack dump on contint1002.wikimedia.org using SIGUSR1 # [[phab:T419009|T419009]]
=== 2026-03-03 ===
* 23:52 James_F: Zuul: [mediawiki/extensions/WikimediaMessages] Drop MetricsPlatform phan dep
* 23:52 James_F: Zuul: [mediawiki/extensions/WikimediaEvents] Drop MetricsPlatform phan dep
=== 2026-03-02 ===
* 22:13 James_F: Zuul: Enforce PHP 8.4 in MW extensions and skins for development branch, for [[phab:T386108|T386108]]
* 14:05 James_F: Zuul: [mediawiki/extensions/MobileFrontend] Add ParserMigration dependency, for [[phab:T415451|T415451]]
* 13:48 James_F: Zuul: […/WikimediaEvents] Drop LoginNotify dependency, now unused, for [[phab:T404334|T404334]]
* 10:16 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/15.8.2/ # [[phab:T418718|T418718]]
=== 2026-02-28 ===
* 21:33 hashar: gerrit: triggering replication to GitHub for all of `mediawiki/skins` # [[phab:T418675|T418675]]
* 21:33 hashar: gerrit: triggering replication to GitHub for all of `mediawiki/extensions` # [[phab:T418675|T418675]]
=== 2026-02-27 ===
* 15:53 dancy: Updating gitlab-cloud-runners (staging and prod) to gitlab-runner 18.9.0.
=== 2026-02-26 ===
* 20:16 James_F: Zuul: Provide a custom, high-priority pipeline just for puppet compiler [[phab:T414621|T414621]]
* 19:32 James_F: Docker: Bump all the PHPs.
* 13:40 hashar: Deployed Jenkins job https://integration.wikimedia.org/ci/job/wikibase-selenium/ # [[phab:T287582|T287582]]
* 00:13 dduvall: forcing replacement of buildkitd helm release in gitlab-cloud-runner prod cluster due to dependency on removed k8s secret ([[phab:T416260|T416260]])
=== 2026-02-25 ===
* 23:50 dduvall: deploying https://gitlab.wikimedia.org/repos/releng/gitlab-cloud-runner/-/merge_requests/552 to gitlab-cloud-runner production cluster ([[phab:T416260|T416260]])
* 14:07 James_F: Zuul: [mediawiki/extensions/CommunityRequests] Add TemplateData dependency, for [[phab:T401638|T401638]]
* 00:08 jeena: no-op testing updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/95
=== 2026-02-24 ===
* 15:55 brennen: devtools: test deploy phab/phorge to test instance ([[phab:T418256|T418256]])
=== 2026-02-23 ===
* 23:07 jeena: Updated development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/92
* 22:43 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/92
* 22:12 bd808: Unblock 191.80.192.0/18 ([[phab:T418132|T418132]])
* 20:26 hashar: Deleted "replication-upstream" Grafana dashboard in favor of a copy/new "replication" one. https://grafana.wikimedia.org/d/RFLS1GsWk/replication-upstream , replaced it by https://grafana.wikimedia.org/d/d4a4da73-c27f-4ce6-a9e5-ab84dd7a4ebb/replication
* 16:29 James_F: Zuul: [3d2png] Add basic Node CI at version 20
=== 2026-02-20 ===
* 21:47 bd808: Unblock 168.184.84.0/24 ([[phab:T418020|T418020]])
* 17:13 bd808: Unblock 122.187.64.0/18 ([[phab:T417964|T417964]])
* 14:35 James_F: Zuul: [mediawiki/extensions/Monstranto] Move out of Wikimedia prod section
=== 2026-02-19 ===
* 18:34 bd808: Unblock 181.98.0.0/16 ([[phab:T417890|T417890]])
* 17:21 James_F: Zuul: [mediawiki/extensions/WikimediaEvents] Add AbuseFilter as a dependency, for [[phab:T417799|T417799]]
* 13:22 hashar: Reloaded Zuul to archive the Cergen repository {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1240688 {{!}} [[phab:T417887|T417887]]
=== 2026-02-18 ===
* 20:17 jeena: Updating development images on contint primary for [[phab:T415922|T415922]]
* 19:44 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240360
* 18:40 bd808: Unblock 46.59.0.0/17 ([[phab:T417747|T417747]])
* 17:05 hashar: Regenerating Jenkins jobs with JJB based on https://gerrit.wikimedia.org/r/c/integration/config/+/1240254/
* 17:04 hashar: Added EXT_DEPENDENCIES to Quibble Jenkins jobs parameters so we can manually trigger them from the Web UI using a different set of deps # https://gerrit.wikimedia.org/r/c/integration/config/+/1240254/
* 16:30 hashar: Triggered https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/ with empty Zuul parameters introduced by https://gerrit.wikimedia.org/r/1240333 {{!}} https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/4893/console
* 15:43 James_F: Zuul: [mediawiki/extensions/ReadingLists] Add EventBus dependency for [[phab:T417706|T417706]]
* 12:15 hashar: zuul-1001.zuul3.eqiad1.wikimedia.cloud: added keepalive=20 to the scheduler Gerrit driver and restarted scheduler container # [[phab:T417497|T417497]]
* 06:58 jeena: Updating development images on contint primary for [[phab:T415922|T415922]]
=== 2026-02-17 ===
* 23:37 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240081
* 23:20 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240078
* 15:58 brennen: deployed latest phab/phorge wmf/stable to devtools test instance ([[phab:T417657|T417657]])
* 09:01 hashar: Reloaded Zuul to enable php 8.5 testing on utfnormal, php-session-serializer, wikipeg, mediawiki/libs/Dodo, mediawiki/libs/UUID, testing-access-wrapper and translatewiki # [[phab:T406326|T406326]]
=== 2026-02-16 ===
* 15:27 hashar: Manually cleaned some old workspaces on integration-agent-docker-1042
=== 2026-02-12 ===
* 20:07 James_F: Zuul: Enable PHP 8.5 jobs for most MW libraries, for [[phab:T406326|T406326]]
* 19:33 James_F: Docker: [php83] Re-build with upstream's new 8.3.30 release and cascade
* 19:31 James_F: Zuul: Add PHP 8.5 CI job to various things noted as blocked by Phan, for [[phab:T410941|T410941]], [[phab:T406326|T406326]]
* 16:35 Krinkle: Disable publishing noise on tasks from repos Bcp47, clover-diff, ScopedCallback, and IDLeDOM. Ref [[phab:T143162|T143162]]
* 15:53 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/87
* 11:21 James_F: Zuul: [mediawiki/libs/shellbox] Add direct Phan job, for [[phab:T416064|T416064]]
=== 2026-02-10 ===
* 20:16 dancy: Rebooted k3s.catalyst-dev (it was unresponsive, but the reboot hasn't helped)
=== 2026-02-09 ===
* 21:58 James_F: Zuul: [mediawiki/tools/phan] Add PHP 8.5 CI job, for [[phab:T410941|T410941]]
* 19:46 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1238006 [[phab:T415680|T415680]]
* 11:51 James_F: Zuul: [mediawiki/extensions/ReadingLists] Drop MetricsPlatform dependency, for [[phab:T414435|T414435]]
=== 2026-02-05 ===
* 17:58 James_F: Zuul: […/WikimediaCustomizations] Add six new dependencies for [[phab:T404334|T404334]]
* 15:35 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1237254
* 15:18 James_F: Zuul: […/OATHAuth] Add dependency and phan dependency on CentralAuth
=== 2026-02-04 ===
* 12:54 James_F: Zuul: [mediawiki/extensions/Petition] Add CLDR dependency
* 10:03 hashar: Restarted Jenkins on releases2003.codfw.wmnet
=== 2026-02-02 ===
* 21:17 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1234926 "re-enable master jobs for some BlueSpice repos - [[phab:T403196|T403196]]"
* 21:05 bd808: Unblock 85.146.0.0/17 ([[phab:T416079|T416079]])
* 19:47 James_F: Zuul: […/WikimediaCustomizations] Add cldr phan dependency, for [[phab:T404334|T404334]]
* 17:33 bd808: Unblock 188.188.0.0/15 ([[phab:T416095|T416095]])
* 17:26 bd808: Unblock 85.94.84.0/22 ([[phab:T416105|T416105]])
* 17:09 bd808: Unblock 94.234.0.0/16 ([[phab:T416165|T416165]])
* 16:51 dancy: Update gitlab-runners to alpine-v18.6.6 ([[phab:T415214|T415214]])
* 16:27 bd808: Unblock 47.231.208.0/21 ([[phab:T416010|T416010]])
* 11:39 James_F: Zuul: […/WikimediaCustomizations] Add five new phan dependencies, for [[phab:T404334|T404334]]
* 09:45 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 58532, 58557
=== 2026-01-31 ===
* 21:49 James_F: Deleted Jenkins's job entry for castor-save-workspace-cache {{Gerrit|6193776}} and this seems to have unstuck things for [[phab:T416078|T416078]]?
* 21:45 James_F: Running `sudo systemctl restart jenkins` on contint for [[phab:T416078|T416078]]
* 21:44 James_F: Fighting [[phab:T416078|T416078]], took integration-castor-5 offline, disconnected, sshed in to kill threads, then reconnected; no change in aspect.
* 19:03 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1235380
=== 2026-01-28 ===
* 21:26 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/WebAuthn # [[phab:T415832|T415832]]
* 21:11 bd808: Unblock 181.160.0.0/15 & 186.40.128.0/17 ([[phab:T415820|T415820]])
* 17:01 bd808: Unblock 102.182.0.0/16 ([[phab:T415782|T415782]])
=== 2026-01-27 ===
* 16:45 James_F: Zuul: Switch skin-quibble template with identical extension-quibble, for [[phab:T402398|T402398]]
* 16:18 James_F: Zuul: [ArticleGuidance] mention it will be in production
* 15:55 James_F: Docker: [quibble-bullseye] Update to Quibble 1.15.0
* 15:12 James_F: Docker: [quibble-coverage] Pass PHPUnit config location explicitly, for [[phab:T395470|T395470]]
* 09:18 hashar: integration: on integration-castor05, deleted caches for old MediaWiki branches
* 09:15 hashar: integration: on pkgbuilder instances, removed Buster cow images, aptcache and hooks. `sudo cumin --force -p 0 'name:pkgbuilder' 'rm -fR /srv/pbuilder/<nowiki>{</nowiki>base-buster-amd64.cow,hooks/buster,aptcache/buster-amd64<nowiki>}</nowiki>'` # [[phab:T397209|T397209]]
* 09:14 hashar: integration: cleaned up old workspaces under /srv/jenkins/workspace
=== 2026-01-26 ===
* 23:27 bd808: Unblock 66.130.0.0/15 ([[phab:T415596|T415596]])
* 22:52 bd808: Unblock 45.16.0.0/12 ([[phab:T415467|T415467]])
* 14:46 hashar: gerrit: changed `operations/software/permissions` project type from `CODE` to `PERMISSIONS` by pointing `HEAD` to `refs/meta/config`
=== 2026-01-22 ===
* 17:36 James_F: Docker: [quibble-coverage] Stop using legacy PHPUnit entrypoint ([[phab:T395470|T395470]]) & Stop excluding Dump/ParserFuzz/Stub groups ([[phab:T415230|T415230]])
* 15:11 James_F: Zuul: [mediawiki/extensions/Math] Add a standalone job, for [[phab:T415230|T415230]]
=== 2026-01-20 ===
* 20:38 bd808: Cherry picked https://gerrit.wikimedia.org/r/c/operations/puppet/+/1229186 ([[phab:T415113|T415113]])
* 19:05 bd808: Rebooted deployment-cache-text08 to see if the mystery haproxy startup failure would go away ([[phab:T415100|T415100]])
* 18:50 bd808: Unblock 152.7.0.0/16 ([[phab:T415100|T415100]])
=== 2026-01-17 ===
* 23:32 ori: beta-scap with `php_l10n: true` completed successfully: https://integration.wikimedia.org/ci/view/Beta/job/beta-scap-sync-world/241466/console. PHP l10n files generated. Reverted local change to scap.cfg.
* 23:26 ori: Temporarily set `php_l10n: true` on deployment-deploy04:/etc/scap.cfg to see if next scap succeeds.
=== 2026-01-16 ===
* 16:33 dancy: Deleting deployment-mx03.deployment-prep ([[phab:T412975|T412975]])
=== 2026-01-15 ===
* 14:50 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/ArticleSummaries/ # [[phab:T413232|T413232]]
=== 2026-01-14 ===
* 17:14 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1226907
* 16:27 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1226893
* 15:57 bd808: Unblock 190.60.63.0/24 ([[phab:T414541|T414541]])
=== 2026-01-13 ===
* 15:04 James_F: Zuul: Make quibble-for-mediawiki-core-vendor-mysql-php84 voting, for [[phab:T386108|T386108]]
=== 2026-01-12 ===
* 21:33 zabe: zabe@deployment-mwmaint03:~$ foreachwiki migrateLinksTable.php --table imagelinks # [[phab:T413668|T413668]]
* 21:06 bd808: Unblock 66.81.168.0/21 ([[phab:T414303|T414303]])
* 17:42 dancy: Turned off instance deployment-prep.deployment-mx03
* 11:44 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 46331, 46344
=== 2026-01-10 ===
* 21:48 taavi: reload zuul for https://gerrit.wikimedia.org/r/1224782
* 00:25 bd808: Unblock 91.160.0.0/12 ([[phab:T414190|T414190]])
=== 2026-01-09 ===
* 17:33 thcipriani: re-enabling beta update jobs after test bad extension-list [[phab:T411516|T411516]]
* 17:09 thcipriani: disabling beta update jobs to test bad extension-list [[phab:T411516|T411516]])
=== 2026-01-08 ===
* 21:30 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1224815 [[phab:T414136|T414136]]
* 18:24 bd808: Unblock 89.80.0.0/12 ([[phab:T414113|T414113]])
* 15:55 dancy: Upgrading gitlab-runner to v18.5.0 on gitlab-cloud-runners. ([[phab:T414053|T414053]])
=== 2026-01-07 ===
* 23:17 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1082574 https://gerrit.wikimedia.org/r/1224157 https://gerrit.wikimedia.org/r/1224159
* 23:12 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/896311 [[phab:T27482|T27482]]
* 23:06 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1224218
* 17:34 James_F: Zuul: Add new extensions: IssueTrackerLinks, PreviewLinks, and WikiRAG
* 17:34 James_F: Zuul: [labs/tools/heritage] Point to the task to drop 8.1 testing
* 15:09 James_F: Zuul: [labs/tools/heritage] Add testing in PHP 8.2+, not just PHP 8.1
* 15:03 James_F: Zuul: Even for extension-broken, don't offer PHP 8.1 testing
* 15:02 James_F: Zuul: Move quibble experimental sqlite/postgres tests to PHP 8.3
=== 2026-01-06 ===
* 16:57 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1223690 [[phab:T411814|T411814]]
* 16:16 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1223189 [[phab:T411814|T411814]]
* 00:30 bd808: Unblock 85.134.128.0/17 ([[phab:T413755|T413755]])
* 00:02 bd808: Unblock 89.166.128.0/17 ([[phab:T413702|T413702]])
=== 2026-01-05 ===
* 23:57 bd808: Unblock 185.233.104.0/22 ([[phab:T413472|T413472]])
* 23:51 bd808: Unblock 45.62.112.0/21 ([[phab:T413079|T413079]])
* 23:44 bd808: Unblock 85.134.200.0/21 ([[phab:T413067|T413067]])
* 19:03 dancy: Updated buildkitd to v0.26.3 in gitlab-cloud-runners
* 14:27 taavi: reload zuul for {{Gerrit|1223191}}
* 13:57 James_F: Zuul: [mediawiki/php/wmerrors] Enable PHP 8.5 testing, for [[phab:T410921|T410921]]
=== 2026-01-03 ===
* 17:59 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1222709 https://gerrit.wikimedia.org/r/1220388 https://gerrit.wikimedia.org/r/1219140
=== 2026-01-02 ===
* 17:10 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1222597
=== 2026-01-01 ===
* 02:34 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1221644
<noinclude>'''Server Admin Log''' logged from {{IRC|wikimedia-releng}} for [[Nova Resource:Deployment-prep|Beta Cluster]], [[mw:Continuous integration|Continuous integration]] and various other Release Engineering projects.</noinclude>
{{SAL-archives/Release Engineering}}
<noinclude>[[Category:SAL]]</noinclude>
qhkzli7n6sftrgys2elkz69htne0p4b
SRE/Service Operations/Ownership
0
448299
2456766
2431842
2026-09-11T12:57:08Z
MLechvien-WMF
49197
2456766
wikitext
text/x-wiki
{{Service_Operations/Navigation}}{{Warn
| content = The list below is deprecated starting from FY26-27, please refer to [https://www.mediawiki.org/wiki/Wikimedia_Production/Service_Catalog Service Catalog]
}}
{| class="wikitable" width=100%
!Service/Procedure
!Description
|-
|[[Kubernetes/Clusters|WikiKube Kubernetes Cluster]]
|Kubernetes (often abbreviated k8s) is an open-source system for automating deployment, and management of applications running in containers.
|-
|[[Application servers|Mediawiki servers]]
|The Application servers (or app servers) are the several hundred Apache servers that run the MediaWiki backend software (written in PHP).
|-
|[[Memcached for MediaWiki]]
|There are two logical pools of memcached servers for MediaWiki. There are critical for performance of all sites and used extensively
|-
|[[Redis|Redis Misc]]
|Redis is used in Wikimedia production for:
changeprop (role::redis::misc)
As a cache and queue backend in ORES
Receiver of sampled profile data from PHP, as part as the sampling/profiling pipeline (Arc Lamp).
|-
|[[Shellbox]]
|Shellbox is a library for remote command execution, and a server for secure command execution. It was primarily implemented to sandbox lilypond (used by the Score extension) and provide a way for MediaWiki to utilize external binaries without needing them to be in the same container. Shellbox relies on Kubernetes (and Linux containers/namespaces) to provide isolation and resource limits for external commands.
|-
|[[Switch_Datacenter|Datacenter Switchover]]
|A datacenter switchover (from eqiad to codfw, or vice-versa) comprises switching over multiple different components, some of which can happen independently and many of which need to happen in lockstep. This page documents all the steps needed to switch over from a master datacenter to another one, broken up by component. SRE Service Operations maintains the process and software necessary to run the switchover.
|-
|[[SLO|Service Level Objectives]]
|Service Level Objective (SLO) and Service Level Indicators (SLI)
|-
|[[Kafka#main (eqiad and codfw)|Kafka-main]]
|kafka-main is the low-volume, critical production services cluster. Talk to us before starting to send events there. kafka-main is currently used directly by [[Event_Platform/EventGate]] and change-propagation.
|}
q91u3t4ujeyc2kgieli92askeck7y2q
Map of database maintenance
0
449160
2456810
2456699
2026-09-12T00:02:18Z
Dexbot
30554
Bot: Updating the report
2456810
wikitext
text/x-wiki
{{/Header}}
== Today (2026-09-12) ==
== Yesterday (2026-09-11) ==
== Last seven days ==
{| class="wikitable"
|+ eqiad
|-
! Section !! Work
|-
| s4 ||
* [[phab:T404715|Setup x4 section (T404715)]] (marostegui)
* [[phab:T437108|Move production reads for commons link tables to x4 (T437108)]] (ladsgroup)
|-
|}
{| class="wikitable"
|+ codfw
|-
! Section !! Work
|-
| s4 || [[phab:T404715|Setup x4 section (T404715)]] (marostegui)
|-
| s5 || [[phab:T437188|Switchover s5 master (db2213 -> db2157) (T437188)]] (marostegui)
|-
| x4 || [[phab:T404715|Setup x4 section (T404715)]] (marostegui)
|-
|}
[[Category:MariaDB]]
9ek9v38i4fmx9ahdejogaaphgfofmyf
Data Platform/Systems/Ceph
0
452006
2456781
2422560
2026-09-11T14:50:52Z
BTullis (WMF)
25295
Removed an incorrect statement about CephFS.
2456781
wikitext
text/x-wiki
The [[Data Platform Engineering]] team manages a pair of [[Ceph]] clusters for two primary purposes:
# To consider the use of Ceph and its S3 compatible interface as a replacement for (or addition to) [[Data Engineering/Systems/Cluster/Hadoop|Hadoop]]'s HDFS file system
# To provide block and file storage capability to workloads running on the [[Kubernetes/Clusters#dse-k8s|dse-k8s]] Kubernetes cluster.
A number of specific use cases have already been implemented, such as:
* [[Postgres]] using the [[Data Platform/Systems/PostgreSQL|cloudnativepg operator]], where primary data is stored on RBD and backups are stored on S3
* [[Data Platform/Systems/Airflow|Airflow]]:
** Task log storage is on S3
** Kerberos credential cache is on cephfs
** DAGs are distributed using cephfs
* [[Event Platform/Stream Processing/Flink|Flink]] checkpoints are stored on S3
== Project Status ==
The cluster is in a production state. We manage five servers in eqiad. There is a smaller cluster of three servers in codfw.
Here is the original [https://docs.google.com/document/d/1dhAlABcM08zMcw9u01qwukhnw2bf6jQ9rKsRkuRRjdQ/view design document] for the project. The Phabricator epic ticket is for the eqiad cluster deployment is: [[phab:T324660|T324660]]
== Software Components ==
The following sections give a brief explanation of each of the four principal software components involved in our [https://docs.ceph.com/en/reef/architecture/#arch-ceph-storage-cluster Ceph Storage Cluster], with a reference to how we run them.
Please refer to [https://docs.ceph.com/en/reef/architecture/ Ceph Architecture] for a more in-depth explanation of each component.
=== Monitor Daemons ===
[https://docs.ceph.com/en/reef/rados/configuration/mon-config-ref/ Ceph monitor daemons] are responsible for maintaining a [https://docs.ceph.com/en/reef/architecture/#architecture-cluster-map cluster map], which keeps an up-to-date record of the cluster topology and the location of data objects.
The Paxos algorithm is used to ensure that a quorum of servers are in agreement about the contents of the cluster map.
We run a monitor daemon on all five of our Ceph servers in eqiad, so our quorum is three active mon servers.
=== Manager Daemons ===
[https://docs.ceph.com/en/reef/mgr/#ceph-manager-daemon Ceph manager daemons] run alongside monitor daemons in order to provide additional monitoring and management interfaces. Only one manager daemon is ever active, but there may be several standby servers and ceph itself manages the election of a new mgr daemon to the active role. In our eqiad cluster we have five manager daemons, since we run one alongside each monitor daemon.
=== Object Storage Daemons ===
A Ceph cluster will contain many [https://docs.ceph.com/en/reef/rados/configuration/storage-devices/ Object Storage Daemons] (referred to as OSDs). There is usually a 1:1 relationship between an OSD and a physical storage device, such as a hard drive (HDD) or a solid-state drive (SSD). These OSDs may started and stopped individually, in order to bring storage devices into and out of service.
Our clusters use the [https://docs.ceph.com/en/reef/rados/configuration/storage-devices/#bluestore Bluestore] specification of OSD, rather than the deprecated [https://docs.ceph.com/en/reef/rados/configuration/storage-devices/#filestore Filestore] specification. This means that each OSD maintains its own RocksDB database of object metatadata, complete with ''write ahead log'' (WAL). The physical location of the WAL and the <code>block.db</code> database file can be specified independently of the OSD backing store, in order to optimise performance and availability. Please see https://docs.ceph.com/en/reef/rados/configuration/bluestore-config-ref/ for a more comprehensive explanation of Bluestore.
In our eqiad cluster we have 20 OSDs per host, for a total of 100 OSDs in the cluster. 60 of these are backed by hard drives and 40 by SSDs. We use a high-performance NVMe drive to host the WAL and RocksDB databases of the HDD backed OSDs.
=== Rados Gateway Daemons ===
The rados gateway daemons (referred to as <code>radosgw</code>) are different from the <code>mon</code>, <code>mgr</code>, and <code>osd</code> daemons in that they are not an internal component of the Ceph cluster. They are ''clients'' of the cluster.
They serve an HTTP interface that enables the [https://docs.ceph.com/en/reef/radosgw/#object-gateway Ceph Object Gateway], which is the S3 and Swift compatible interface to the storage services.
On our clusters we currently run a rados gateway on each of our <code>cephosd</code> hosts. The hostname of our S3/Swift interfaces are
* https://rgw.eqiad.dpe.anycast.wmnet
* https://rgw.codfw.dpe.anycast.wmnet
=== Metadata Server Daemons ===
The [https://docs.ceph.com/en/reef/glossary/#term-MDS Metadata Server (MDS)] component supports [https://docs.ceph.com/en/reef/cephfs/ Cephfs], which is a shared file system using POSIX semantics.
We run an MDS daemon on each of our hosts. In eqiad we have three Ceph file systems, which means that we always have three MDS daemons active and two standby.
== Cluster Architecture ==
At present we run a co-located configuration on our cluster in eqiad comprising five hosts, each of which is in a 2U enclosure that is optimised for storage.
The server names are: <code>cephosd100[1-5].eqiad.wmnet</code> and they are identical to each other in terms of hardware.
Each server contains two shelves, each containing twelve hot-swappable 3.5" drive bays. At the rear of the chassis are two hot-swappable bays, which contain drives that are used for the operating system.
Each host currently has a 10 Gbps network connection to its switch. If we start to hit this throughput ceiling, we can request to increase this to 25 Gbps by changing the optics.
=== Storage Device Configuration ===
Each of the five hosts has the following primary storage devices:
[[File:Ceph_Server_Storage_Diagram.png|alt=Diagram of the storage devices within each ceph server|left|thumb]]
The specific storage devices are as follows:
{| class="wikitable"
|+
!Count
!Capacity
!Technology
!Make/Model
!Total Capacity
!Use Case
|-
|12
|18 TB
|HDD
|Seagate Exos X18 nearline SAS
|216 TB
|Cold tier
|-
|8
|3.8 TB
|SSD
|Kioxia RM6 mixed-use
|30.4 TB
|Hot tier
|}
That makes the raw capacity of the five-node cluster:
* Cold tier: '''1.08 PB'''
* Hot tier: '''152 TB'''
In order to increase the performance of the cold tier, which is backed by hard drives, we employ an NVMe device as a cache device. This is partitioned such that each of the HDD based OSD daemons can store its [https://docs.ceph.com/en/latest/rados/configuration/storage-devices/#bluestore Bluestore] database and journal on it, increasing performance of these systems considerably.
=== Ceph Software Configuration ===
Each of our hosts runs the following Ceph services:
* 1 monitor daemon (<code>ceph-mon</code>)
* 1 manager daemon (<code>ceph-mgr</code>)
* 1 metadata service daemon (<code>ceph-mds</code>)
* 20 object storage daemons (<code>ceph-osd</code>)
* 1 crash monitoring daemon (<code>ceph-crash</code>)
* 1 rados gateway daemon (<code>radosgw</code>)
<syntaxhighlight lang="bash">
btullis@cephosd1004:~$ pstree -T|egrep 'ceph|radosgw'
|-ceph-crash
|-ceph-mds
|-ceph-mgr
|-ceph-mon
|-20*[ceph-osd]
|-radosgw
</syntaxhighlight>This cluster uses Ceph packages that are distributed by the upstream project at: https://docs.ceph.com/en/latest/install/get-packages/#apt and are integrated with [[reprepro]].
[[Puppet]] is used to configure all of the daemons on this cluster.
=== Additional Services ===
Alongside the Ceph daemons, each of the hosts in the cluster runs the following components:
* An instance of [[envoy]] which is operating as a TLS termination endpoint for the local <code>radosgw</code> service.
* An instance of bird, which is providing support for the [[anycast]] load-balancing that we use for the <code>radosgw</code> service.
== Cluster Configuration ==
=== Configuration Methods ===
The cluster configuration is partly managed with puppet and partly by-hand.
==== Puppet Configuration ====
We intend to migrate more of the configuration into puppet over time. One thing to bear in mind is that we share the [https://github.com/wikimedia/operations-puppet/tree/production/modules/ceph ceph puppet module] with the [[WMCS]] team, as they also use it for their clusters. However, we apply different puppet profiles ([https://github.com/wikimedia/operations-puppet/tree/production/modules/profile/manifests/ceph '''ceph'''] vs [https://github.com/wikimedia/operations-puppet/tree/production/modules/profile/manifests/cloudceph '''cloudceph''']) to our clusters, in order to allow for some variance in the configuration style. Be aware of this potential impact on the WMS team when modifying puppet. In addition to this, be aware that there is a [https://github.com/wikimedia/operations-puppet/tree/production/modules/profile/manifests/cephadm '''cephadm''' profile] and a [https://github.com/wikimedia/operations-puppet/tree/production/modules/cephadm '''cephadm''' module] in puppet. These are '''not''' relevant to the configuration of the DPE Ceph cluster, as they are only in use for the new [[Ceph#Clusters|apus cluster]] that is managed by the [[Data persistence|Data Persistence]] team.
Elements of the cluster configuration that are managed by puppet:
* Ceph package installation.
* The <code>/etc/ceph/ceph.conf</code> configuration file.
* All daemons (<code>mon</code>, <code>mgr</code>, <code>osd</code>, <code>crash</code>, <code>radosgw</code>)
* Preparation and activation of the hardware storage devices beneath each of the <code>osd</code> daemons.
* All [https://docs.ceph.com/en/reef/rados/operations/user-management/ Ceph client users] (Note that this is a different concept from a Ceph object storage user, which are not yet managed with puppet).
==== Manual Configuration ====
The following are currently managed by hand on the cluster.
* Pool creation
* Associating pools to applications
* CRUSH maps and rules, which affect data placement
* Maintenance flags, such as <code>noout</code>
* <code>radosgw</code> user management, for S3 and Swift access (Created {{Phabricator/en|T374531}} to address this.)
=== CRUSH Rules ===
CRUSH is an acronym for ''Controlled Replication Under Scalable Hashing''. It is the algorithm that determines where in the storage cluster any item of data should reside, including attributes such as the number of replicas of that item and/or any parity information that would allow the item to be reconstructed in the event of any loss of a storage device.
For more detailed information on CRUSH, please refer to https://docs.ceph.com/en/reef/rados/operations/crush-map/
CRUSH ''rules'' are used to generate CRUSH ''maps.'' The idea behind a CRUSH map is that the Ceph monitor (aka mon) servers load the CRUSH maps into memory and enables clients to locate data within the cluster. When a client requests data from a Ceph cluster, the mon responds with the location of the data including on which osd(s) the data resides. The idea is to avoid a network bottleneck, since the mon does not proxy the data itself. Clients communicate directly with the osd processes when reading and writing data. This is analagous to the way in which HDFS namenodes provide a metadata service for clients to communicate with HDFS datanodes.
We currently have the following two CRUSH rules in place on the DPE Ceph cluster.<syntaxhighlight lang="bash">
btullis@cephosd1005:~$ sudo ceph osd crush rule ls
hdd
ssd
btullis@cephosd1004:~$ sudo ceph osd crush rule dump
[
{
"rule_id": 1,
"rule_name": "hdd",
"type": 1,
"steps": [
{
"op": "take",
"item": -4,
"item_name": "default~hdd"
},
{
"op": "chooseleaf_firstn",
"num": 0,
"type": "host"
},
{
"op": "emit"
}
]
},
{
"rule_id": 2,
"rule_name": "ssd",
"type": 1,
"steps": [
{
"op": "take",
"item": -6,
"item_name": "default~ssd"
},
{
"op": "chooseleaf_firstn",
"num": 0,
"type": "host"
},
{
"op": "emit"
}
]
}
]
</syntaxhighlight>The <code>hdd</code> and <code>ssd</code> rules are for using replicated pools, selecting for the corresponding device classes.
We have configured our cluster with buckets for <code>row</code> and <code>rack</code> awareness, so the CRUSH algorithm is aware of the host placement within the rows.<syntaxhighlight lang="bash">
btullis@cephosd1005:~$ sudo ceph osd crush tree|grep -v 'osd\.'
ID CLASS WEIGHT TYPE NAME
-1 1149.93103 root default
-19 689.95862 row eqiad-e
-21 229.98621 rack e1
-3 229.98621 host cephosd1001
-22 229.98621 rack e2
-7 229.98621 host cephosd1002
-23 229.98621 rack e3
-10 229.98621 host cephosd1003
-20 459.97241 row eqiad-f
-24 229.98621 rack f1
-13 229.98621 host cephosd1004
-25 229.98621 rack f2
-16 229.98621 host cephosd1005
</syntaxhighlight>The osd objects (currently numbered 0-99) are then assigned to the host objects, so our data is always distributed between hosts, rows, and racks.
=== Pools ===
In a Ceph cluster, a [https://docs.ceph.com/en/reef/rados/operations/pools/ pool] is a logical partitioning of objects. Pools are [https://docs.ceph.com/en/reef/rados/operations/pools/#associating-a-pool-with-an-application associated with an application], specifically the <code>mgr</code>, <code>rbd</code>, <code>radosgw</code>, or <code>cephfs</code> applications.
We can list the pools with the <code>ceph osd pool ls</code> or <code>ceph osd lspools</code> commands.<syntaxhighlight lang="bash">
btullis@cephosd1004:~$ sudo ceph osd lspools
2 .mgr
7 dse-k8s-csi-ssd
8 .rgw.root
9 eqiad.rgw.log
10 eqiad.rgw.control
11 eqiad.rgw.meta
12 eqiad.rgw.buckets.index
13 eqiad.rgw.buckets.data
14 eqiad.rgw.buckets.non-ec
</syntaxhighlight>All pool names beginning with a <code>.</code> are reserved for use internally by the cluster, so do not attempt to modify these pools. In our case we have the .mgr pool which is in use by the
In our case, we have created the <code>dse-k8s-csi-ssd</code> pool for use with the <code>rbd</code> application and the [[Data Platform/Systems/Ceph#Kubernetes Integration|kubernetes Integration]]. It is backed by SSDs and is a replicated pool with 3 replicas.
All of those pools with rgw in their name are related to the radosgw application and underpin the [[Data Platform/Systems/Ceph#Object Storage|S3/Swift object storage]] capabilities.
=== File Systems ===
We have three file systems in eqiad. These are <code>dpe</code>, <code>dumps</code>, and <code>home</code>.<syntaxhighlight lang="bash">
btullis@cephosd1001:~$ sudo ceph fs ls
name: dpe, metadata pool: cephfs.dpe.meta, data pools: [cephfs.dpe.data-ssd ]
name: dumps, metadata pool: cephfs.dumps.meta, data pools: [cephfs.dumps.data ]
name: home, metadata pool: cephfs.home.meta, data pools: [cephfs.home.data ]
</syntaxhighlight>The dpe file system is something of a general purpose file system that is backed by the SSD pools. This is used in Airflow for the DAGs and kerberos credential caches.
The others are backed by the HDDs.
All metadata is stored on SSDs.
==== MDS Configuration ====
We have increased the value for the [https://docs.ceph.com/en/reef/cephfs/cache-configuration/#confval-mds_cache_memory_limit mds_cache_memory_limit] value from the default of 4Gi to 8Gi. See {{Phabricator/en|T401094}} for details.
== Kubernetes Integration ==
We have enabled [https://docs.ceph.com/en/latest/rbd/rbd-kubernetes/ Kubernetes block devices] on the dse-k8s cluster, by means of the [https://github.com/ceph/ceph-csi/ Ceph-CSI] (Container Storage Interface) project.
== Object Storage ==
We have enabled the [https://docs.ceph.com/en/latest/radosgw/ Ceph Object Gateway] (radosgw) in order to provide S3 and Swift compatible APIs and object storage.
[[Category:Data platform]]
[[Category:Data platform systems]]
== User Management ==
=== Creating a ceph user ===
SREs can execute the following command from one of the ceph hosts (such as ''cephosd1001'').
<syntaxhighlight lang="bash">
sudo radosgw-admin user create --uid=example --display-name="An Example"
</syntaxhighlight>
You can check your work with the below command (also includes expected output).
<syntaxhighlight lang="bash">
sudo radosgw-admin user info --uid=example
{
"user_id": "example",
"display_name": "example",
"email": "",
"suspended": 0,
"max_buckets": 1000,
"subusers": [],
"keys": [
{
"user": "example",
"access_key": "[...]",
"secret_key": "[...]"
}
],
"swift_keys": [],
"caps": [],
"op_mask": "read, write, delete",
"default_placement": "",
"default_storage_class": "",
"placement_tags": [],
"bucket_quota": {
"enabled": true,
"check_on_raw": false,
"max_size": -1,
"max_size_kb": 0,
"max_objects": -1
},
"user_quota": {
"enabled": false,
"check_on_raw": false,
"max_size": 4398046511104,
"max_size_kb": 4294967296,
"max_objects": -1
},
"temp_url_keys": [],
"type": "rgw",
"mfa_ids": []
}
</syntaxhighlight>
=== Setting quotas ===
<syntaxhighlight lang="bash">
sudo radosgw-admin quota set --quota-scope=user --uid=research --max-size=4T
</syntaxhighlight>
=== Checking quotas ===
=== Accessing the S3 endpoint ===
==== Sample .s3cfg for s3cmd tool ====
== Upgrading ==
Please see: [[Data Platform/Systems/Ceph/Upgrading]]
== Troubleshooting ==
See [[Data Platform/Systems/Ceph/Troubleshooting]]
qeeqz5jyspdmxcb8pbq3qhnv7ztml2d
SRE/Service Operations/Documentation/Reboots
0
453463
2456805
2452871
2026-09-11T20:34:38Z
RLazarus (WMF)
15215
2456805
wikitext
text/x-wiki
There are times when we need to reboot our fleet. Here are some notes to help us go through it like
== Datastores ==
=== [[Memcached for MediaWiki | Memcache]] cluster ===
Servers of this cluster need no depooling of any sort, and we have a pool of servers to pick up the traffic of any unavailable server. The cached data in the memecache cluster are crucial for our latency, so it is highly recommended to '''reboot 1 server at a time per DC''', with a sleep time of 15'-20' minimum between reboots, allowing the server to warm up.
* https://grafana-rw.wikimedia.org/d/000000316/memcache?orgId=1
==== Cookbook: <code>sre.memcached.roll-reboot-restart</code> ====
Use the [[Spicerack/Cookbooks|cookbook]] <code>sre.memcached.roll-reboot-restart</code> to perform rolling reboots or daemon restarts on memcached hosts.
The cookbook automatically:
* Verifies the gutter pool is healthy before touching the main pool (as failover capacity)
* Confirms memcached is active on each host after the operation
===== Options =====
Standard cookbook options are used, with the addition of <code>--min-uptime</code>, so to provide some easy way to resume operations.
* '''Currently the <code> --query </code> is broken, please only use aliases'''
* <code>--min-uptime</code>: Only include hosts whose memcached service has been running at least this long (e.g. <code>7d</code>, <code>24h</code>, <code>604800</code>)
===== Examples =====
Rolling reboot of eqiad, one host at a time with 15-minute sleep (recommended defaults):
<pre>
cookbook sre.memcached.roll-reboot-restart --reason "Debian reboots" \
--alias memcached-eqiad --batchsize 1 --grace-sleep 900 reboot
</pre>
Restart only hosts whose memcached has been running for at least 10 days (e.g. to resume interrupted rolling restarts):
<pre>
cookbook sre.memcached.roll-reboot-restart --reason "Resume reloads" \
--alias memcached-eqiad --min-uptime 10d --batchsize 2 restart_daemons
</pre>
=== [[Redis]] lock ===
Redis-lock hosts can be rebooted one after the other, MediaWiki can handle loss of one server at a time. There is no replication or complex mechanisms, just check the service is back up.
Use the <code>sre.hosts.reboot</code> cookbook.
=== [[Redis]] misc ===
Redis hosts can be rebooted one after the other, taking care of waiting for replication to be back up after the reboot.
Dashboards: [[https://grafana.wikimedia.org/d/000000174/redis?from=now-3h&orgId=1&timezone=utc&to=now&var-instance=$__all&var-job=redis_misc&var-site=$__all redis-misc on Grafana]]
'''Please do not restart these servers during deployments because docker-registry uses redis'''
==== Service Upgrade ====
'''Upgrading redis-server'''
A simple <code>apt-get install -y redis-server</code> is enough, but it will '''not''' trigger a restart, as we are using one systemd service per port, thus you will see the following message
redis-server.service is a disabled or a static unit not running, not starting it.`
'''Restarting instances'''
Always restart '''replicas first''', then '''primaries second'''.
Restart all instances one by one with a 5-second delay:
<syntaxhighlight lang="bash">
rdb1001:# for port in 6378 6379 6380 6381 6382; do echo "Restarting $port..."; systemctl restart redis-instance-tcp_$port; sleep 5; done
</syntaxhighlight>
==== Server reboot ====
Run the <code>sre.hosts.reboot</code> cookbook, taking care to reboot the replica first, and to wait for the replication status to be ok before moving to the primary.
==== Common checks ====
'''Checking replication status'''
To check the replication status of all instances:
<syntaxhighlight lang="bash">
rdb1001:# for port in 6378 6379 6380 6381 6382; do echo "--- $port ---"; redis-cli -p $port -a $(grep requirepass /etc/redis/tcp_$port.conf | awk '{print $2}') --no-auth-warning info replication | grep -E "master_link_status|master_last_io_seconds_ago|master_sync_in_progress"; echo "-----"; done
</syntaxhighlight>
Example output:
<syntaxhighlight lang="text">
--- 6379 ---
master_link_status:down
master_last_io_seconds_ago:-1
master_sync_in_progress:1
-----
--- 6380 ---
master_link_status:up
master_last_io_seconds_ago:0
master_sync_in_progress:0
-----
</syntaxhighlight>
A healthy, caught-up replica should show:
* <code>master_link_status:up</code>
* <code>master_last_io_seconds_ago</code>: expect a low value (0-10)
* <code>master_sync_in_progress:0</code>
'''Full resync after restart'''
After a restart, the replica will perform a '''full resync''' since it has no cached master state:
<syntaxhighlight lang="text">
Master replied to PING, replication can continue...
Partial resynchronization not possible (no cached master)
Full resync from master: 7fe8b694f6afc193b57187a779dfd037c9884ee3:0
</syntaxhighlight>
During this process, <code>master_last_io_seconds_ago:-1</code> indicates the replica has not yet connected or is still performing a full sync.
'''Note:''' In the case of a '''primary''' redis server restart with a large dataset, it may take a while for replicas to reconnect, as the primary will need to load the full dataset into memory first.
==== docker-registry ====
<code>docker-registry</code> doesn't use HA for redis, so the reboot of its redis nodes will cause a short unavailability. '''Avoid restarting these nodes during deployments''', as the unavailability of the docker-registry redis could cause deployment failures.
----
=== Etcd ===
== Kubernetes ==
=== Production ===
==== Workers ====
<code>
sudo cookbook -d sre.k8s.reboot-nodes --batchsize 15 --k8s-cluster wikikube-codfw --reason "Reason" --alias wikikube-worker-codfw --minimal-cordon
</code>
<code>
sudo cookbook -d sre.k8s.reboot-nodes --batchsize 15 --k8s-cluster wikikube-eqiad --reason "Reason" --alias wikikube-worker-eqiad --minimal-cordon
</code>
==== Control plane ====
<code>
sudo cookbook -d sre.k8s.reboot-nodes --batchsize 1 --k8s-cluster wikikube-eqiad --reason "Reason" --alias wikikube-master-eqiad --minimal-cordon
</code>
==== Dragonfly supernodes====
Lock scap to avoid overloading the registry servers with a big deployment
<code>deploy1003:~$ scap lock --all "Dragonfly supernodes reboot"</code>
Then reboot using the <code>sre.hosts.reboot-single</code> cookbook
=== Staging ===
== PoolCounter ==
{{See|"PoolCounter" redirects here. For the software's documentation, see [[mw:PoolCounter]].}}
Each poolcounter server should be removed from <code>mediawiki-config/wmf-config/ProductionServices.php</code> with a gerrit patch before reboot, then that patch should be deployed to production using <code>scap backport</code>.
After the reboot, add the server back and remove the next server in the same patch, <code>scap backport</code> it, and repeat until all poolcounter servers are rebooted and back in <code>mediawiki-config/wmf-config/ProductionServices.php</code>
[[Thumbor]] uses poolcounter as well, but will fail open if its poolcounter server is down. You can forgo swapping the servers out in thumbor's <code>deployment-charts/helmfile.d/services/thumbor/values-{eqiad,codfw}.yaml</code> files if the interruption is short.
== Chartmuseum ==
Chartmuseum is addressed through <code>helm-charts.discovery.wmnet</code>. It is backed by one VM in each datacentre.
<syntaxhighlight lang=shell-session>
# Depool codfw
sudo confctl --object-type discovery select 'dnsdisc=helm-charts.*,name=codfw' set/pooled=false
# Reboot codfw
sudo cookbook sre.hosts.reboot-single -r "May 2025 Reboots" chartmuseum2001.codfw.wmnet
# Repool codfw
sudo confctl --object-type discovery select 'dnsdisc=helm-charts.*,name=codfw' set/pooled=true
# Rinse and repeat for eqiad
sudo confctl --object-type discovery select 'dnsdisc=helm-charts.*,name=eqiad' set/pooled=false
sudo cookbook sre.hosts.reboot-single -r "May 2025 Reboots" chartmuseum1001.eqiad.wmnet
sudo confctl --object-type discovery select 'dnsdisc=helm-charts.*,name=eqiad' set/pooled=true
</syntaxhighlight>
[[Category:SRE Service Operations]]
5btb63khg4hstnaowmgumj0qum9lvvp
Ircstream
0
455726
2456765
2433342
2026-09-11T11:59:55Z
MMuhlenhoff (WMF)
5268
/* Bots still using the legacy setup */ Update Phab reference for CVNBot
2456765
wikitext
text/x-wiki
{{Navigation Wikimedia infrastructure|expand=infrafun}}
'''Ircstream''' is an IRC service for broadcasting recent changes events from Wikimedia wikis. It was created in 2024 as new backend for [[irc.wikimedia.org]] and is developed by [[User:Faidon Liambotis|Faidon]] at https://github.com/paravoid/ircstream
== Service ==
The service is currently hosted on irc1003.wikimedia.org and irc2003.wikimedia.org. Events are broadcasted from [[MediaWiki On Kubernetes]] pods (and a handful of remaining baremetal MediaWiki [[app servers]]) via UDP packets sent to port 9390. The ''ircstream.wikimedia.org'' CNAME points to the currently active server.
== Administration ==
=== Adding/removing hosts ===
Two steps are needed to add ircstream servers that MediaWiki uses:
# First, update the network policies for MediaWiki pods, e.g. https://gerrit.wikimedia.org/r/c/operations/deployment-charts/+/1076730
# Second, update MediaWiki configuration, e.g. https://gerrit.wikimedia.org/r/c/operations/mediawiki-config/+/1050261
When removing servers, follow these steps in reverse.
=== Updating the Debian package ===
The upstream repository includes a debian/ directory, so it's a simple rebuild of what's shipped there.
== Use ==
We now have a a vastly superior [[EventStreams]] service providing machine-readable push notifications over HTTP in JSON format at https://stream.wikimedia.org/v2/stream/recentchange, but until the key consumers of the IRC recent changes feed have migrated, this old service remains vital. No new services should use this service.
== Bots still using the legacy setup ==
{| class="wikitable"
|+Overview of legacy setup bots
|-
!scope="col"| Bot name
!scope="col"| Description
!scope="col"| Homepage
!scope="col"| Confirmed to actively consume the IRC feed
!scope="col"| If so, bug for eventstream port
|-
!scope="row"| ClueBot
| ClueBot NG is an anti-vandalism bot that tries to detect and revert vandalism quickly and automatically || https://github.com/cluebotng/bot/ and https://github.com/cluebotng/bot/ || Fixed to use Eventstreams
|https://github.com/cluebotng/bot/issues/58
|-
!scope="row"| COIBot
| COIBot is a bot that tries to track edits that are made by users who may have a conflict of interest || https://meta.wikimedia.org/wiki/User:COIBot/COIBot || ||
|-
!scope="row"| CVNBot
| CVNBot (previously known as SWMTBot) is the main IRC bot software used by the Countervandalism Network to interpret recent changes streams and provide a feed on IRC for patrollers to work from || [[m:CVNBot]] ([https://gerrit.wikimedia.org/g/labs/countervandalism/CVNBot gerrit]) ([[phab:tag/cvnbot/|phab]]) || Yes, ~30 bots with some number of active users || [[phab:T437537]]
|-
!scope="row"| EarwigBot
| EarwigBot is a Python bot that edits Wikipedia and interacts over IRC || https://github.com/earwig/earwigbot || ||https://github.com/earwig/earwigbot/issues/85
|-
!scope="row"| EstadoEdita
| Bot which publishes Wikipedia edits from the State of Chile to Twitter, apparently defunct since no posts after May 2023 || https://github.com/jsajuria/estadoedita/ || seems defunct || n/a
|-
!scope="row"| EyeInTheSkyBot
| This is a stalk bot for edits on Wikipedia. It uses the RC IRC feed on irc.wikimedia.org to look for edits, and reports stalked edits in ##eyeinthesky on irc.libera.chat || https://github.com/stwalkerster/eyeinthesky || yes || https://github.com/stwalkerster/eyeinthesky/issues/61
|-
!scope="row"| FinlandEdit
| Bot which publishes Wikipedia edits from different Finnish Government IPs to Twitter (defunct since July 2023) and Mastodon || https://github.com/duukkis/finlandedits || || https://github.com/edsu/anon/issues/144
|-
!scope="row"| KrdBot
| The bot is written in Perl and based on the MediaWiki::Bot-Module. It runs continuously automatically or manually assisted at specified intervals, depending on tasks (see below). It normally does not more than one edit per minute on the same task || https://commons.wikimedia.org/wiki/User:Krdbot || Fixed to no longer needed IRC||https://commons.wikimedia.org/wiki/User_talk:Krd#c-Ladsgroup-20260702170000-Using_IRC_streams
|-
!scope="row"| OverheidsEdits
| This Twitter bot warns you when the Dutch government anonymously edits Wikipedia articles. Seems defunct, no posts to Twitter since July 2023 and the listed Mastodon URL is dead too || || seems defunct ||
|-
!scope="row"| Huggle
| Huggle is a diff browser intended for dealing with vandalism and other unconstructive edits on Wikimedia projects, written in C++ using the Qt framework. Huggle is a desktop tool, not a bot, but it uses IRC as one of three RC sources (alongside the often-unreliable [[XmlRcs]] and RecentChanges polling). IRC has been recently preferred by users due to its reliability (compared to XmlRcs) and speed (compared to the polling API). || [[m:Huggle]] ([https://github.com/huggle/huggle3-qt-lx GitHub]) ([[phab:tag/huggle/|Phab]]) || Fixed to no longer rely on IRC || https://github.com/huggle/huggle3-qt-lx/issues/391
|-
!scope="row"| WikiMon
| Watch the RecentChanges IRC feed with Python. Support for various WikiMedia projects and languages. At the moment, WikiMon's primary usage pattern is broadcasting changes over WebSocket || https://github.com/hatnote/wikimon || Yes, actively used by http://listen.hatnote.com/ || none
|-
!scope="row"| Pastilles
| Watch the RecentChanges IRC feed with || https://fr.wikipedia.org/wiki/Utilisateur:PastilleBot || yes || none
|}
== See also ==
* Source code: https://github.com/paravoid/ircstream (Upstream homepage)
* List of bots on meta: [[m:IRC/Bots]]
* Operational dashboard: https://grafana.wikimedia.org/d/eb101795-c69e-4b9c-b848-f042d604f234/ircstream?orgId=1
[[Category:IRC]]
[[Category:SRE Infrastructure Foundations]]
9y6scqlh27wdomz5giylbykzbt91u8g
Data Platform Engineering/Ops week/Analytics weekly train
0
459290
2456782
2456531
2026-09-11T15:16:20Z
XCollazo-WMF
31846
2456782
wikitext
text/x-wiki
...🚂🚃🚃🚃🚃🚃🚃🚃🚃🚃🚃🚃🚃🚃
== Analytics deployment train ==
☑️ Only add here stuff that has been merged.
☑️ Link the task and the Gerrit patch.
☑️ List the systems that need deploying, jar versions that need bump-ups, and jobs that need restarting, if there are any.
Extra points if you include what to run and where to run it (e.g. stat1007, an-coord1001...).
☑️ Do you have a way of checking the deployment has been successful?
☑️ Don't move stuff to "ready to deploy" in the kanban unless it's documented here.
☑️ Check [[Data Engineering/Ops week#The Data Engineering deployment train 🚂|Data_Engineering/Ops_week#The_Data_Engineering_deployment_train_'''🚂''']] for a '''pointer about Wikistats, as well as links for various types of deployments.'''
☑️ To see the old log, go to [[etherpad:p/analytics-weekly-train/timeslider#59750|https://etherpad.wikimedia.org/p/analytics-weekly-train/timeslider#59747]].
'''Now use the log below.''' Eventually we could have some sub-pages or templates to streamline this.
=== YYYY-MM-DD NEXT TUESDAY TRAIN (REPLACE THIS AFTER DEPLOY) ===
Deployer: XXX
Refinery:
* 1335995: Fix pageview_per_editor_per_page_daily inflating view counts | <nowiki>https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1335995</nowiki>
* T437608 please merge hql/fr_tech/centralnoticebannerhistory_events.hql https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1338956
* T431256 Add X-Provenance to `wmf.webrequest` | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1308104 Operations needed, ping @joal :)
Refinery-source:
* 1338971: MWHistoryDeltaWriter: align fields with mediawiki_history | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1338971
=== 2026-09-01 Tuesday ===
Deployer: [[User:JAllemandou (WMF)|JAllemandou (WMF)]] ([[User talk:JAllemandou (WMF)|talk]]) 15:23, 1 September 2026 (UTC)
Refinery:
* 1332814: File Export: HQL for point-in-time snapshot of mediawiki content | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1332814
* 1328608: webrequest_sampled: Add is_thumb_generated field | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1328608
=== 2026-08-25 Tuesday ===
Deployer: xcollazo
Refinery:
* 1325932: Update user_agents_info with Cloudflare Radar | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1325932
Refinery-source:
* 1327582: refinery-job: Add HiveToJdbc job to mirror Hive tables to PostgreSQL | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1327582
=== 2026-08-12 Wednesday ===
Deployer: [[m:User:Mforns_(WMF)|User:Mforns (WMF)]]
Refinery:
* 1319913: HQLs to create and import Cloudflare bot directory | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1319913
=== 2026-07-29 Wednesday ===
Deployer: [[User:AKhatun (WMF)]]
Refinery:
* 1318667: Update media jobs for thumbnail change | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1318667
* 1319107: Provide sqoopable_dbnames for tables that were dropped in some wikis | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1319107
* 1319028: spark: Add logback config to quiet DataHub/OpenLineage lineage agent | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1319028
Refinery source:
=== 2026-07-22 Wednesday ===
Deployer: [[User:AMastilovic-WMF]]
Refinery:
* 1311398: Add creation statement for daily sqooped tables cu_log and logging. | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1311398
* 1310561: Add revert_risk_wikidata prediction to event sanitization allowlist. | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1310561
Refinery source:
=== 2026-07-14 Tuesday: ===
Deployer: [[User:APizzata-WMF]]
Refinery:
* 1309640: Ops Week fix: missing comma | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1309640
* 1309654: Add new fields to webrequest_sampled | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1309654
Refinery source:
* 1308685: Remove any reference to the windowing of 90 Days. | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1308685
=== 2026-07-10 Friday ===
Deployer: [[User:JMonton-WMF]]
Refinery:
* 1308190 | Add hql/fr_tech/banner_activity_minutely.hql: https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1308190
* 1306491 | DDL for wmf_traffic.user_agents_info table: https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1306491
* 1308121| druid: add two new dimensions to the webrequest sampled's indexing: https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1308121
=== Tuesday, July 07, 2026 ===
Deployer: [[User:JMonton-WMF]]
Refinery:
* Update sqoop and <code>wmf.mediawiki_filerevision</code> table: https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1308087 (requires operations on the table: drop, recreate + repair)
=== Tuesday, June 30, 2026 ===
Deployer: dr0ptp4kt
Status: '''Deployed'''
Notes: Although the deployment appeared to succeed, this did not seem to address sqoop errors concerning apiportalwiki.
Refinery:
* Pageview allowlist changes
** https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1305158
** https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1305162
** https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1305156
**
* Pageview allowlist and sqoop changes
** https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1305980
* Add new tables to Sqoop
** https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1295064
** https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1295069
Refinery source:
* 1305625: Update MWH - fail if duplicated revisions | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1305625
* 1305633: Update Inc-MWH splitting SQL into smaller files | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1305633
* 1305674: Distribute and sort snapshot rows on write | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1305674<br />
'''<big>Tuesday, June 23, 2026</big>'''
Deployer: [[User:SNwachukwu (WMF)|Sandra]]
Refinery:
* [[phab:T427532|T427532 Sqoop globalimagelinks and filerevision]] patches
* 1298321: Update pageview allowlist https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1298321 https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1302183
* 1299532: Fix unique-devices Cassandra key for www. projects | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1299532
* Update Incremental MWH schema for readability (no-op) https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1302825 and https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1303466
'''<big>Friday, June 19th, 2026</big>'''
Deployer: [[User:JAllemandou (WMF)|JAllemandou (WMF)]] ([[User talk:JAllemandou (WMF)|talk]])
Refinery source:
* 1298333: Add Iceberg WAP branching to MWHistoryDeltaWriter and MWHistorySnapshotMerger | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1298333
* 1298378: Fix control_map timestamps: use ISO-8601 UTC format (yyyy-MM-dd'T'HH:mm:ss.SSS'Z') uniformly | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1298378
* 1300876: MWHistoryDeltaWriter: fix CAST format bug and stale comment | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1300876
* 1301381: incremental mediawiki history: Add ingestion pipeline diagram | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1301381
=== Monday, June 8th, 2026 ===
Deployer: [[User:JAllemandou (WMF)|JAllemandou (WMF)]] ([[User talk:JAllemandou (WMF)|talk]]) 10:02, 8 June 2026 (UTC)
Refinery:
* 1294816: Move the deletion for content table to private | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1294816
* 1297686: User-Agent compliance and API requests refinement HQLs | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1297686
* 1287959: Add DDL for mediawiki_history_incremental_v1 Iceberg table | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1287959
Refinery source:
* 1290064: Ingest revision_visibility_change to populate revision_deleted_parts | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1290064
* 1290172: Add event_entity='page' to MWHistoryDeltaWriter (MERGE 5) | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1290172
* 1290945: Add event_entity='user' to MWHistoryDeltaWriter (MERGE 6) | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1290945
* 1293836: Daily revert detection: align with monthly DenormalizedRevisionsBuilder | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1293836
* 1296665: Fix MWHistorySnapshotMerger: DELETE+INSERT replaces MERGE, add page/user reconcile | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1296665
* 1297743: Add event_user_is_cross_wiki, page_is_deleted, revision_is_deleted_by_page_deletion, user_central_id | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1297743
* 1297115: Fix MWH revert algorithm | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1297115
=== Wednesday, May 27, 2026 ===
Deployer: [[User:JAllemandou (WMF)|JAllemandou (WMF)]] ([[User talk:JAllemandou (WMF)|talk]]) 19:26, 27 May 2026 (UTC)
Refinery:
* 1288444: Determine API requests in webrequest refine job | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1288444
* 1290844: Optimize Cassandra load mediarequest_top_files to avoid OOM. | <nowiki>https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1290844</nowiki>
* 1285337: Move mediawiki_content to private folder | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1285337
Source:
* 1293659: change mediawiki_content to private | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1293659
=== Monday, May 18, 2026 ===
Deployer: [[User:TChin (WMF)]]
Refinery Source:
* 1286385: Add event_user_is_cross_wiki to wmf.mediawiki_history | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1286385
* 1285904: Add event_log_id to wmf.mediawiki_history | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1285904
* 1286481: Upgrade graphframes to 0.11.0 from Maven Central, drop Archiva repos | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1286481
* 1286397: Refactor MediawikiEvent.fromRow to use named column access | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1286397
* 1286989: Remove wmf-analytics-old-uploads Archiva repository | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1286989
* 1287508: Add Sanitizer to clean up wprov value of x-analytics. | <nowiki>https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1287508</nowiki>
Refinery:
please deploy Refinery after deploying Refinery Source above.
* 1285903: Add event_log_id to mediawiki_history DDL | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1285903
* 1285527: querypage: Add UncategorizedImages.hql | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1285527
* [[gerrit:c/analytics/refinery/+/1283749|expand_event_sanitized_analytics_allowlist: Add revertrisk-multilingual predictions to allowlist. (1283749)]]
* 1286383: Add event_user_is_cross_wiki to mediawiki_history DDL | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1286383
* 1287443: Add mediawiki_page_html_feature_counts_change_v1 to allowlist for event_sanitized | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1287443
* 1287909: Use SanitizeXAnalyticsWprovUDF to normalize x_analytics[wprov] values | <nowiki>https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1287909</nowiki>
=== '''Thursday, May 07, 2026''' ===
Deployer: Aisha
Refinery:
* 1277789: querypage: Add WantedCategories.hql | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1277789
* 1278746: Remove mediawiki_revision_score from sanitization main allowlist | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1278746
* 1279483: Remove SearchSatisfaction from sanitization analytics allowlist | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1279483
* 1279651: mediarequest_hourly: use file/filetypes as media_classification ground truth | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1279651
=== '''Wednesday, April 29, 2026''' ===
Deployer: Antonio
Refinery:
* move hql script from fundraising to fr_tech | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1260793
* 1267966: querypage: MostCategories: Include all content namespaces | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1267966
* 1276836: querypage: Add UnusedTemplates.hql | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1276836
* 1275815: Update webrequest validation algorithm | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1275815
* 1275982: Remove DesktopWebUIActionsTracking, MobileWebUIActionsTracking, ReadingDepth from sanitization allowlist | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1275982<br />
=== '''Thursday, March 25, 2026''' ===
Deployer: Aisha and Sandra
Refinery:
* Add abstract.wikipedia to pageview allowlist | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1256413
* Changes to mapper-weight for centralauth_localuser | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1256302
* Move bot detection pipeline into new repo | <nowiki>https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1237928</nowiki>
=== Thursday, March 10, 2026 ===
By mforns
Refinery:
* Add kai.wikipedia to the pageview allowlist
https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1249328
DONE (sync'ed by hand)
Airflow:
* Artifact cleaning: remove outdated refinery job artifacts | https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/2030#d6baf582888b568d8a7bcb95316bd03cbefa9853 | '''Note this change might cause some backward compatibility issues and we would need to monitor the DAGs closely after deployment.'''
DONE
=== 2026-03-11 ===
Deployer: joal
Refinery-source:
* Update ProduceCanaryEvents job https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1249982 + https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1250016
=== 2026-03-05 (special Thursday post cleanups) ===
Deployer: dr0ptp4kt (with Marcel and Sandra)
Refinery:
* Adapt imagelinks pipeline and consumers for imagelink normalization | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1239200
* No-op: Fix druid banner_activity data prep job | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1240253
Airflow:
* After refinery deployment. Pass mediawiki_private_linktarget_table to commons impact metrics dag | https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/2026
=== 2026-02-18 ===
Deployer: joal
Refinery:
* Use names in banner activity GROUP BY - https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1239232
* Add first_campaign_status_code for banner activity - https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1238821
=== 2026-02-10 ===
Deployer: xcollazo
Refinery:
* 1235830: MediawikiDumper: fix filenames to include end revision when covering a single page. | <nowiki>https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1235830</nowiki>
* 1236347: Migrate cu_changes table to use cuua_text in new cu_usergent table. | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1236347
=== 2026-02-01 ===
Deployer: Joseph
Refinery:
* 1233834: Remove mediawiki_wikitext_* from refinery-drop-mediawiki-snapshots | https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1233834
** Minor non-urgent patch. No need to release if just this patch.
* Update pageview project allowlist - https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1235201
* HQL for druid webrequest_sampled ingestion https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1235740
Airflow:
* Load webrequest_sampled in druid hourly https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1967
=== 2026-01-21 ===
Deployer: Antonio/Joseph
*Refinery
**Update pingback HQL code for new PHP and MediaWiki versions https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1222506
**Update pageview allowlist
***https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1225039
***https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1229091
***https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1229095
***https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1229507
**Update event _sanitized allowlist https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1207489
*Airflow
**Update pingback MediaWiki and PHP versions to include new values https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1909
***We need the refinery deployment done first
***After deploying this, from Cindy: Once the patches are merged, the weekly queries will need to be re-run starting from the beginning of May 2025. Xcollazo is happy to do this part after we deploy. Just ping Xcollazo.
=== 2025-12-03 ===
Deployer: Antoine
'''Refinery:'''
* {{PhabT|T409584}} Add JA3N User-Agent queries https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1212214 and https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1213488 and https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1213522 (no need to do anything else!)
=== 2025-11-18 ===
Deployer: Marcel and Javier
'''Refinery:'''
* {{PhabT|405039}} - Add HQL for edit_per_editor_per_page_daily and pageview_per_editor_per_page_daily https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1196892
DONE
=== 2025-11-12 ===
Deployer: Joal
'''Refinery-source:'''
* {{PhabT|406531}} - Add new referral sources to pageview data https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1203389
* {{PhabT|408178}} - Remove mediawiki.wikistories_* santization allowlist entries https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1202718
* '''[[phab:T407239|T407239]]''' '''-''' Fix Duplicate Pageview metrics records in data quality tables. | <nowiki>https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1203129</nowiki>
* [[phab:T406000|T406000 Adapt mediawiki_history to the removal of mediawiki revision.rev_sha10]] ({{Gerrit|1202334}})
* 1203124: Fix bug MW Dumper in which vertical bars ( `|` ) were not being honored. | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1203124
** After refine-source release, we should:
*** merge https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1795 that will pick up this fix on the File Export DAGs
*** wait until merge request makes it to main Airflow instance
*** delete DagProperties at https://airflow.wikimedia.org/variable/edit/372 , so that the auto-regenerated one points to new jar
*** resume the following DAGs, which have been cleared and are ready to go:
**** https://airflow.wikimedia.org/dags/mw_content_xml_export_current_mid_month/grid
**** https://airflow.wikimedia.org/dags/mw_content_xml_export_current_monthly/grid
**** https://airflow.wikimedia.org/dags/mw_content_xml_export_history_monthly/grid
'''Airflow:'''
* {{PhabT|406531}} - Add new referral sources to pageview data - https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1796
* {{PhabT|409470}} - Fix mediawiki_history_dumps failure - https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1797
* https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1795 (see above in refinery-source section)
=== 2025-11-05 ===
Deployer: Joseph
Refinery Source:
* 1199485: Add Data quality check for Pageview Human-Bot ratio anomaly | [[gerrit:c/aalytics/refinery/source/+/1199485|https://gerrit.wikimedia.org/r/c/aalytics/refinery/source/+/1199485]]
* {{PhabT|406531}} - Add new referral sources to pageview data https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1198313
* Mediawiki-History Bug fix: https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1202191
Airflow:
* [[phab:T407239|T407239]] - Add Dag to run daily Human to Bot page views ratio check https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1776 This MR should be deployed after refinery source is deployed. It needs refinery-job jar v0.3.7
* {{PhabT|406531}} - Add new referral sources to pageview data https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1780 This MR should be deployed after refinery source is deployed. It needs refinery-hive jar v0.3.7
'''<big>2025-10-29</big>'''
deployer: Sandra
Refinery Source:
* 1198080: Fix various bugs on MW Dumper code. | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1198080
* 1198152: Add utility to create SHA256 fingerprints of the files of a particular HDFS folder. | https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1198152
=== 2025-10-22 ===
To-be deployer: Aleksander
* Refinery Source
** Add user_central_id to the mediawiki_history dataset(s) https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1194951
=== 2025-10-14 ===
To-be deployer: Marcel
*Refinery
**{{PhabT|405533}} - Unique devices data uses non-standard domains for Wikidata, Wikifunctions, and MediaWiki.org<nowiki/> https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1194885 . Note: This task has a pending Ai<nowiki/>rflow patch to be merged/deployed once this one is deployed: [[gitlab:repos/data-engineering/airflow-dags/-/merge_requests/1743|htt]]<nowiki/>[[gitlab:repos/data-engineering/airflow-dags/-/merge_requests/1743|p]][[gitlab:repos/data-engineering/airflow-dags/-/merge_requests/1743|s://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1743]] [DONE]
**{{PhabT|406000}} - Adapt mediawiki_history to the removal of mediawiki revision.rev_sha1 - https://gerrit.wikimedia.org/r/c/analytics/refinery/+/1196716 Nullify sha1 in Sqoop [DONE]
*Refinery Source
**[[phab:T365203|T365203]] - Add check for wikis count to Mediawiki history dat<nowiki/>a quality checks [[gerrit:c/analytics/refinery/source/+/1193440|h]]<nowiki/>[[gerrit:c/analytics/refinery/source/+/1193440|ttps://gerrit.wikim]]<nowiki/>[[gerrit:c/analytics/refinery/source/+/1193440|edia.org/r/c/analytics/refinery/source/+/1193440]] [DONE]
**[[phab:T365203|T365]]<nowiki/>[[phab:T365203|203]] - Bug Fix: Add support for Deequ Metric value Distribution d<nowiki/>ata type [[gerrit:c/analytics/refinery/source/+/1195268|https://gerrit.wikim]]<nowiki/>[[gerrit:c/analytics/refinery/source/+/1195268|edia.org/r/c/analytics/refinery/source/+/1195268]] [DONE]
**{{PhabT|406000}} - Adapt mediawiki_history to the removal of mediawiki revision.rev_sha1 - [[gerrit:c/analytics/refinery/source/+/1196049|https://gerrit.wikime]][[gerrit:c/analytics/refinery/source/+/1196049|dia.org/r/c/analytics/refinery/source/+/1196049]] and [[gerrit:c/analytics/refinery/source/+/1196469|https://gerrit.wikimedia.or]][[gerrit:c/analytics/refinery/source/+/1196469|g/r/c/analytics/refinery/source/+/1196469]]. Note: This patch needs a related Airflow patch: https://gitlab.wikimedia.org/repos/data-engineering/airflow-dags/-/merge_requests/1750. This one also: https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1196485 [DONE]
**{{PhabT|T384945}} Modify code to dump all slots AND {{PabT|T405641}} Adapt MW Content pipelines to the removal of upstream revision.rev_sha1 - https://gerrit.wikimedia.org/r/c/analytics/refinery/source/+/1195330 [DONE]
d3ykt8bd7rgi9noeflg1slb36atvszn
Wikidata Query Service/Migration/Rewrite of GAS/Examples
0
460602
2456795
2454900
2026-09-11T17:46:35Z
AWesterinen-WMF
54189
/* GAS' maxIterations Compared to pathSearch's maxDepth */
2456795
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE { ?predecessor wdt:P40 ?item }
}
}
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
}
}
} GROUP BY ?item
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE { ?item wdt:P131 ?predecessor } # Reverse direction
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?item
}
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } # Reverse direction
}
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?item
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE
{ ?predecessor wdt:P279 ?out }
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?out
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
Why are the results different?
There are indeed 36 results following the subclassOf path (human being a subclassOf ?x) for 5 hops.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion. If no maxIterations limit is set or the value is set to greater than 5, it is rewritten to set the maxDepth to 5.
== Returning the Predecessor for Each Item on the Search Path ==
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
18xl9aoa41fxp7u0gdkxagjsf6c09dj
2456796
2456795
2026-09-11T17:49:15Z
AWesterinen-WMF
54189
/* GAS' maxIterations Compared to pathSearch's maxDepth */
2456796
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE { ?predecessor wdt:P40 ?item }
}
}
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
}
}
} GROUP BY ?item
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE { ?item wdt:P131 ?predecessor } # Reverse direction
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?item
}
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } # Reverse direction
}
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?item
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE
{ ?predecessor wdt:P279 ?out }
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?out
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human being a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion and/or deep depth and breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5.
== Returning the Predecessor for Each Item on the Search Path ==
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
o9mwqsmhlgxc85bjskoqfpb4vr6zwif
2456797
2456796
2026-09-11T17:52:01Z
AWesterinen-WMF
54189
/* Why are the results different? */
2456797
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE { ?predecessor wdt:P40 ?item }
}
}
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
}
}
} GROUP BY ?item
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE { ?item wdt:P131 ?predecessor } # Reverse direction
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?item
}
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT *
WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } # Reverse direction
}
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?item
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE
{ ?predecessor wdt:P279 ?out }
}
}
BIND(( ?edge + 1 ) AS ?d)
}
}
} GROUP BY ?out
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
c9kjicvxs20yd21xtj0wyep6r8yc1ry
2456798
2456797
2026-09-11T18:49:44Z
AWesterinen-WMF
54189
2456798
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P40 ?item } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?item wdt:P131 ?predecessor } } # Reverse direction
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } } # Reverse direction
}
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P279 ?out } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?out # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the number of results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
'''Find all descendants of Elizabeth II (Q9682) up to 5 levels deep and list their parents'''
This should output the children, grandchildren, great-grandchildren, etc. of Elizabeth II. A grandchild's parent should be one of the descendants from the previous expansion/depth.
Blazegraph query:
<syntaxhighlight lang="text">
SELECT ?vertex ?predecessor ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q9682 ;
gas:linkType wdt:P40 ;
gas:maxIterations 5 ;
gas:out ?vertex ;
gas:out1 ?depth ;
gas:out2 ?predecessor }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 118 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?vertex ?predecessor ?depth
WHERE {
{ SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
WHERE {
{ SELECT ?vertex (MIN(?d1) AS ?depth)
WHERE {
{ { SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 1st pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d1)
} # Close the 1st pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d1)
}
} # Close the 1st BIND UNION SERVICE
} GROUP BY ?vertex # Close the inner embedded SELECT's WHERE
} # Close SELECT ?vertex (MIN(?d1) AS ?depth)
{ { SERVICE pathSearch:
{ _:b1 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc__1 ;
pathSearch:edgeColumn ?edge__1
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 2nd pathSearch parameter definition
BIND(( ?edge__1 + 1 ) AS ?d2)
} # Close the 2nd pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d2)
}
} # Close the 2nd BIND UNION SERVICE
FILTER ( ?d2 = ?depth )
} GROUP BY ?vertex ?depth # Close the outer embedded SELECT's WHERE
# Close SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 126 msec
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
a38q0gcnfvpmna26rusruadebt2k978
2456799
2456798
2026-09-11T18:50:17Z
AWesterinen-WMF
54189
/* Base Template (Single Seed, `Forward` Search) */
2456799
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P40 ?item } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?item wdt:P131 ?predecessor } } # Reverse direction
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } } # Reverse direction
}
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P279 ?out } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?out # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the number of results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
'''Find all descendants of Elizabeth II (Q9682) up to 5 levels deep and list their parents'''
This should output the children, grandchildren, great-grandchildren, etc. of Elizabeth II. A grandchild's parent should be one of the descendants from the previous expansion/depth.
Blazegraph query:
<syntaxhighlight lang="text">
SELECT ?vertex ?predecessor ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q9682 ;
gas:linkType wdt:P40 ;
gas:maxIterations 5 ;
gas:out ?vertex ;
gas:out1 ?depth ;
gas:out2 ?predecessor }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 118 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?vertex ?predecessor ?depth
WHERE {
{ SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
WHERE {
{ SELECT ?vertex (MIN(?d1) AS ?depth)
WHERE {
{ { SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 1st pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d1)
} # Close the 1st pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d1)
}
} # Close the 1st BIND UNION SERVICE
} GROUP BY ?vertex # Close the inner embedded SELECT's WHERE
} # Close SELECT ?vertex (MIN(?d1) AS ?depth)
{ { SERVICE pathSearch:
{ _:b1 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc__1 ;
pathSearch:edgeColumn ?edge__1
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 2nd pathSearch parameter definition
BIND(( ?edge__1 + 1 ) AS ?d2)
} # Close the 2nd pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d2)
}
} # Close the 2nd BIND UNION SERVICE
FILTER ( ?d2 = ?depth )
} GROUP BY ?vertex ?depth # Close the outer embedded SELECT's WHERE
# Close SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 126 msec
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
bcnkgxaddny1t07nufqap4xt32nf80l
2456800
2456799
2026-09-11T18:50:45Z
AWesterinen-WMF
54189
/* Reverse Search */
2456800
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P40 ?item } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?item wdt:P131 ?predecessor } } # Reverse direction
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } } # Reverse direction
}
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P279 ?out } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?out # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the number of results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
'''Find all descendants of Elizabeth II (Q9682) up to 5 levels deep and list their parents'''
This should output the children, grandchildren, great-grandchildren, etc. of Elizabeth II. A grandchild's parent should be one of the descendants from the previous expansion/depth.
Blazegraph query:
<syntaxhighlight lang="text">
SELECT ?vertex ?predecessor ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q9682 ;
gas:linkType wdt:P40 ;
gas:maxIterations 5 ;
gas:out ?vertex ;
gas:out1 ?depth ;
gas:out2 ?predecessor }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 118 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?vertex ?predecessor ?depth
WHERE {
{ SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
WHERE {
{ SELECT ?vertex (MIN(?d1) AS ?depth)
WHERE {
{ { SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 1st pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d1)
} # Close the 1st pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d1)
}
} # Close the 1st BIND UNION SERVICE
} GROUP BY ?vertex # Close the inner embedded SELECT's WHERE
} # Close SELECT ?vertex (MIN(?d1) AS ?depth)
{ { SERVICE pathSearch:
{ _:b1 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc__1 ;
pathSearch:edgeColumn ?edge__1
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 2nd pathSearch parameter definition
BIND(( ?edge__1 + 1 ) AS ?d2)
} # Close the 2nd pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d2)
}
} # Close the 2nd BIND UNION SERVICE
FILTER ( ?d2 = ?depth )
} GROUP BY ?vertex ?depth # Close the outer embedded SELECT's WHERE
# Close SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 126 msec
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
9fjt9eg6e5d1cjy4uft7ttntxhoiy3x
2456801
2456800
2026-09-11T18:51:41Z
AWesterinen-WMF
54189
/* Reverse Search */
2456801
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P40 ?item } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?item wdt:P131 ?predecessor } } # Reverse direction
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } } # Reverse direction
}
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="text">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P279 ?out } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?out # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the number of results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
'''Find all descendants of Elizabeth II (Q9682) up to 5 levels deep and list their parents'''
This should output the children, grandchildren, great-grandchildren, etc. of Elizabeth II. A grandchild's parent should be one of the descendants from the previous expansion/depth.
Blazegraph query:
<syntaxhighlight lang="text">
SELECT ?vertex ?predecessor ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q9682 ;
gas:linkType wdt:P40 ;
gas:maxIterations 5 ;
gas:out ?vertex ;
gas:out1 ?depth ;
gas:out2 ?predecessor }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 118 msec
QLever query:
<syntaxhighlight lang="text">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
SELECT ?vertex ?predecessor ?depth
WHERE {
{ SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
WHERE {
{ SELECT ?vertex (MIN(?d1) AS ?depth)
WHERE {
{ { SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 1st pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d1)
} # Close the 1st pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d1)
}
} # Close the 1st BIND UNION SERVICE
} GROUP BY ?vertex # Close the inner embedded SELECT's WHERE
} # Close SELECT ?vertex (MIN(?d1) AS ?depth)
{ { SERVICE pathSearch:
{ _:b1 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc__1 ;
pathSearch:edgeColumn ?edge__1
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 2nd pathSearch parameter definition
BIND(( ?edge__1 + 1 ) AS ?d2)
} # Close the 2nd pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d2)
}
} # Close the 2nd BIND UNION SERVICE
FILTER ( ?d2 = ?depth )
} GROUP BY ?vertex ?depth # Close the outer embedded SELECT's WHERE
# Close SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 126 msec
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
5oykc9ky72bt1avlddms371r2664i08
2456802
2456801
2026-09-11T18:55:41Z
AWesterinen-WMF
54189
/* Examples */
2456802
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P40 ?item } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?item wdt:P131 ?predecessor } } # Reverse direction
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } } # Reverse direction
}
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P279 ?out } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?out # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the number of results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
'''Find all descendants of Elizabeth II (Q9682) up to 5 levels deep and list their parents'''
This should output the children, grandchildren, great-grandchildren, etc. of Elizabeth II. A grandchild's parent should be one of the descendants from the previous expansion/depth.
Blazegraph query:
<syntaxhighlight lang="sparql">
SELECT ?vertex ?predecessor ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q9682 ;
gas:linkType wdt:P40 ;
gas:maxIterations 5 ;
gas:out ?vertex ;
gas:out1 ?depth ;
gas:out2 ?predecessor }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 118 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?vertex ?predecessor ?depth
WHERE {
{ SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
WHERE {
{ SELECT ?vertex (MIN(?d1) AS ?depth)
WHERE {
{ { SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 1st pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d1)
} # Close the 1st pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d1)
}
} # Close the 1st BIND UNION SERVICE
} GROUP BY ?vertex # Close the inner embedded SELECT's WHERE
} # Close SELECT ?vertex (MIN(?d1) AS ?depth)
{ { SERVICE pathSearch:
{ _:b1 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc__1 ;
pathSearch:edgeColumn ?edge__1
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 2nd pathSearch parameter definition
BIND(( ?edge__1 + 1 ) AS ?d2)
} # Close the 2nd pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d2)
}
} # Close the 2nd BIND UNION SERVICE
FILTER ( ?d2 = ?depth )
} GROUP BY ?vertex ?depth # Close the outer embedded SELECT's WHERE
} # Close SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 126 msec
=== Why is this so complex? ===
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
77kfbptuofetl8rvbr5iodsgg21lu0d
2456803
2456802
2026-09-11T19:10:56Z
AWesterinen-WMF
54189
/* Returning the Predecessor for Each Item on the Search Path */
2456803
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P40 ?item } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?item wdt:P131 ?predecessor } } # Reverse direction
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } } # Reverse direction
}
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P279 ?out } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?out # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the number of results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
'''Find all descendants of Elizabeth II (Q9682) up to 5 levels deep and list their parents'''
This should output the children, grandchildren, great-grandchildren, etc. of Elizabeth II. A grandchild's parent should be one of the descendants from the previous expansion/depth.
Blazegraph query:
<syntaxhighlight lang="sparql">
SELECT ?vertex ?predecessor ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q9682 ;
gas:linkType wdt:P40 ;
gas:maxIterations 5 ;
gas:out ?vertex ;
gas:out1 ?depth ;
gas:out2 ?predecessor }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 118 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?vertex ?predecessor ?depth
WHERE {
{ SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
WHERE {
{ SELECT ?vertex (MIN(?d1) AS ?depth)
WHERE {
{ { SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 1st pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d1)
} # Close the 1st pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d1)
}
} # Close the 1st BIND UNION SERVICE
} GROUP BY ?vertex # Close the inner embedded SELECT's WHERE
} # Close SELECT ?vertex (MIN(?d1) AS ?depth)
{ { SERVICE pathSearch:
{ _:b1 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc__1 ;
pathSearch:edgeColumn ?edge__1
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 2nd pathSearch parameter definition
BIND(( ?edge__1 + 1 ) AS ?d2)
} # Close the 2nd pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d2)
}
} # Close the 2nd BIND UNION SERVICE
FILTER ( ?d2 = ?depth )
} GROUP BY ?vertex ?depth # Close the outer embedded SELECT's WHERE
} # Close SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 126 msec
=== Why is this so complex? ===
The complexity arises due to the semantic mismatch between Blazegraph's and QLever' search primitives. The query itself ("descendants up to 5 levels, with parents") is simple, but the two available engines' computations are '''not''' equivalent.
Blazegraph's gas:service BFS is an single-pass, breadth-first algorithm which visits each vertex once, at the depth it was first discovered. It records whichever predecessor discovered it — thereby enforcing "shortest path, one parent per vertex" semantics. The seed itself (depth 0) is natively included.
QLever's closest analogue to this is pathSearch:allPaths and is a depth-first algorithm. It follows '''every''' path from the seed up to maxDepth. This is more expensive than "visiting each vertex once." For example, if a grandchild is reachable via two different children (e.g., through two different marriages/lines), allPaths will return the grandchild twice, at two different depths, each with different predecessors. There is no concept of "first thread wins". The query itself has to reduce this down to 1 row per vertex.
The reduction is what adds complexity. It needs three separate steps:
* MIN(depth) per vertex to aggregate over all paths to find each vertex's shortest depth (since a vertex can appear at multiple depths)
* Re-run the pathSearch twice, filtering down to only the rows sitting exactly at the minimum depth
** It is not possible to obtain the "minimum" and "row that produced it" with one pass
** SAMPLE(predecessor) to arbitrarily pick one predecessor among any that tie at the minimum depth (since there is no "first" thread to naturally break the tie)
Further complicating matters, the seed's depth-0 row (wd:Q9682 itself, with no predecessor) is not output by pathSearch:allPaths at all. Only paths with ≥1 edge are produced. For this reason, depth 0's results have to be manually injected via the UNION/BIND.
Another item to note is that the predecessor '''may be different''' when comparing the Blazegraph and QLever results. But, this will only occur if a vertex can be encountered via multiple paths traversed in the search. Admittedly, this is rare, but it can occur. It occurs because Blazegraph's predecessor is the first encountered, whereas QLever's is sampled from all possible predecessors in the path.
== Using Multiple Seeds (gas:in Parameters) ==
== Using a Target (gas:target Parameter) ==
[[Category:WDQS]]
1ljaqtna9zb7moycehl7g64btm9dgnq
2456804
2456803
2026-09-11T20:25:03Z
AWesterinen-WMF
54189
2456804
wikitext
text/x-wiki
= Overview =
A description of the Blazegraph GAS (gather-apply-scatter) WDQS service is provided on the Wikitech page,
[https://wikitech.wikimedia.org/wiki/Blazegraph_Migration:_Rewrite_of_GAS Rewriting the Blazegraph GAS Service]. The examples below illustrate how different GAS service requests can be rewritten to execute on the new QLever backend.
The following examples are provided:
* Base Template (Single Seed, `Forward` Search, No Target, No Predecessor Information)
* Excluding the Depth 0 Seed
* Reverse Search
* Undirected Search
* GAS' maxIterations Compared to pathSearch's maxDepth
* Getting a Predecessor
* Using Multiple Seeds (gas:in Parameters)
* Using a Target (gas:target Parameter)
Note that in the examples, one will see reference to either the GAS BFS (breadth-first) or SSSP (shortest-path) algorithms. Both translate to the same rewritten syntax on QLever since SSSP regresses to BFS (there are no weighted edge paths in Wikidata).
Each Blazegraph example is followed by its rewritten (QLever) syntax along with a comparison of the output results. Execution times are also provided. These are as reported from the current WDQS endpoint (https://query.wikidata.org/) and from the public QLever endpoint for Wikidata (provided by the University of Freiburg, https://qlever.dev/wikidata). Note the the Freiburg endpoint is used since it is the only external QLever endpoint currently available (until WDQS is fully migrated).
Also, it is important to note that your query results may show differences across the two engines (where no differences are reported here). That is because the QLever endpoint may not be using the same Wikidata image (full set of RDF triples) as the current WDQS. WDQS is updated in near real-time.
The rewritten results are the output from the [https://gitlab.wikimedia.org/repos/wikidata-platform/wikidata-query-rewriter rewrite tool], but whitespace has been removed to condense the text - for ease of readability.
= Examples =
== Base Template (Single Seed, `Forward` Search) ==
'''The children of Genghis Khan''' <br><br>
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?itemLabel ?pic ?linkTo
WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q720 ;
gas:traversalDirection "Forward" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P40 .
}
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 796 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?pic ?linkTo
WHERE {
# Blazegraph breadth-first search must be translated to QLever depth-first search
# Need min depth on QLever since it finds an `item` on all paths
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q720 AS ?item) # To output depth=0, the initial seed
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
# A blank node (_:b0) is not needed in the results
# But needed to create a valid triple
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q720 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P40 ?item } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d) # Depth = edge count + 1
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P40 ?linkTo }
OPTIONAL { ?item wdt:P18 ?pic }
}
</syntaxhighlight>
Number of results: 395 results <br>
Execution time: 81 msec
== Reverse Search ==
'''Find all palaces/grand buildings located in any of the administrative territories of Madrid'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q5756 ;
gas:traversalDirection "Reverse" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 4 ;
gas:linkType wdt:P131 .
}
?item wdt:P31/wdt:P279* wd:Q16560 . }
</syntaxhighlight>
Number of results: 188 results <br>
Execution time: 3960 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5756 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5756 ;
pathSearch:maxDepth 4 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?item wdt:P131 ?predecessor } } # Reverse direction
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
?item wdt:P31/(wdt:P279)* wd:Q16560
}
</syntaxhighlight>
Number of results: 108 results <br>
Execution time: 1193 msec
=== Why are the results so different? ===
Actually, the item sets are identical.
Deduplicating both sets of results, Blazegraph has 105 unique ?item values, QLever has 105 unique ?item values, and they are the same items.
The "188 vs 108" gap is entirely duplicate rows, and the duplication is wildly uneven per engine. For example: Q171517 appears 13 times in Blazegraph's output versus 3 times in QLever's; Q3533603 appears 7 times versus 2 times. Every other item beyond those two is a single row in QLever, but as many as 6 rows in Blazegraph.
But, again, why? The duplication comes from ?item wdt:P31/(wdt:P279)* wd:Q16560. This property-path join is appended after the GAS/pathSearch block, and is unchanged by the rewrite.
A transitive * path can be satisfied by more than one distinct derivation chain between the same two nodes, and the SPARQL spec doesn't mandate how many solution rows an engine emits. Blazegraph and QLever count the multiplicity differently.
== Undirected Search ==
'''Find every place within 2 administrative distances of California (in either direction)'''
And for each place, optionally shows what its parent location is.
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?item ?linkTo {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.SSSP" ;
gas:in wd:Q99 ;
gas:traversalDirection "Undirected" ;
gas:out ?item ;
gas:out1 ?depth ;
gas:maxIterations 2 ;
gas:linkType wdt:P131 .
}
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 3236 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?item ?linkTo
WHERE {
{ SELECT ?item (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q99 AS ?item)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q99 ;
pathSearch:maxDepth 1 ;
pathSearch:start ?predecessor ;
pathSearch:end ?item ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE {
{ ?predecessor wdt:P131 ?item } # Forward direction
UNION
{ ?item wdt:P131 ?predecessor } } # Reverse direction
}
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?item # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
OPTIONAL { ?item wdt:P131 ?linkTo }
}
</syntaxhighlight>
Number of results: 71505 results <br>
Execution time: 1972 msec
== GAS' maxIterations Compared to pathSearch's maxDepth ==
'''Find all the superclasses of human (Q5) up to 10 deep'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?out ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" .
gas:program gas:in wd:Q5 .
gas:program gas:out ?out .
gas:program gas:out1 ?depth .
gas:program gas:maxIterations 10 .
gas:program gas:linkType wdt:P279 .
}
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 50 results <br>
Execution time: 485 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?out ?outLabel ?depth WHERE {
{ SELECT ?out (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q5 AS ?out)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q5 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?out ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor wdt:P279 ?out } }
} # Close pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?out # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 36 results <br>
Execution time: 97 msec
=== Why are the number of results different? ===
There are indeed 36 results from Blazegraph following the subclassOf path (human is a subclassOf ?x) for '''5 hops'''.
The original query asked for 10 hops, but the rewritten query changed that to 5. This is done as a protective mechanism to prevent unlimited path expansion (if no limit is set) and/or deep depth and then breadth expansion.
If no maxIterations limit is set on the Blazegraph query, or the value is set to greater than 5, it is rewritten to set QLever's maxDepth to 5. This can be changed in the actual query submitted to QLever.
== Returning the Predecessor for Each Item on the Search Path ==
'''Find all descendants of Elizabeth II (Q9682) up to 5 levels deep and list their parents'''
This should output the children, grandchildren, great-grandchildren, etc. of Elizabeth II. A grandchild's parent should be one of the descendants from the previous expansion/depth.
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?vertex ?predecessor ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q9682 ;
gas:linkType wdt:P40 ;
gas:maxIterations 5 ;
gas:out ?vertex ;
gas:out1 ?depth ;
gas:out2 ?predecessor }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 118 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?vertex ?predecessor ?depth
WHERE {
{ SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
WHERE {
{ SELECT ?vertex (MIN(?d1) AS ?depth)
WHERE {
{ { SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 1st pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d1)
} # Close the 1st pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d1)
}
} # Close the 1st BIND UNION SERVICE
} GROUP BY ?vertex # Close the inner embedded SELECT's WHERE
} # Close SELECT ?vertex (MIN(?d1) AS ?depth)
{ { SERVICE pathSearch:
{ _:b1 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q9682 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor__1 ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc__1 ;
pathSearch:edgeColumn ?edge__1
{ SELECT * WHERE { ?predecessor__1 wdt:P40 ?vertex } }
} # Close 2nd pathSearch parameter definition
BIND(( ?edge__1 + 1 ) AS ?d2)
} # Close the 2nd pathSearch SERVICE
UNION
{ BIND(wd:Q9682 AS ?vertex)
BIND(0 AS ?d2)
}
} # Close the 2nd BIND UNION SERVICE
FILTER ( ?d2 = ?depth )
} GROUP BY ?vertex ?depth # Close the outer embedded SELECT's WHERE
} # Close SELECT ?vertex ?depth (SAMPLE(?predecessor__1) AS ?predecessor)
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 28 results <br>
Execution time: 126 msec
=== Why is this so complex? ===
The complexity arises due to the semantic mismatch between Blazegraph's and QLever' search primitives. The query itself ("descendants up to 5 levels, with parents") is simple, but the two available engines' computations are '''not''' equivalent.
Blazegraph's gas:service BFS is an single-pass, breadth-first algorithm which visits each vertex once, at the depth it was first discovered. It records whichever predecessor discovered it — thereby enforcing "shortest path, one parent per vertex" semantics. The seed itself (depth 0) is natively included.
QLever's closest analogue to this is pathSearch:allPaths and is a depth-first algorithm. It follows '''every''' path from the seed up to maxDepth. This is more expensive than "visiting each vertex once." For example, if a grandchild is reachable via two different children (e.g., through two different marriages/lines), allPaths will return the grandchild twice, at two different depths, each with different predecessors. There is no concept of "first thread wins". The query itself has to reduce this down to 1 row per vertex.
The reduction is what adds complexity. It needs three separate steps:
* MIN(depth) per vertex to aggregate over all paths to find each vertex's shortest depth (since a vertex can appear at multiple depths)
* Re-run the pathSearch twice, filtering down to only the rows sitting exactly at the minimum depth
** It is not possible to obtain the "minimum" and "row that produced it" with one pass
** SAMPLE(predecessor) to arbitrarily pick one predecessor among any that tie at the minimum depth (since there is no "first" thread to naturally break the tie)
Further complicating matters, the seed's depth-0 row (wd:Q9682 itself, with no predecessor) is not output by pathSearch:allPaths at all. Only paths with ≥1 edge are produced. For this reason, depth 0's results have to be manually injected via the UNION/BIND.
Another item to note is that the predecessor '''may be different''' when comparing the Blazegraph and QLever results. But, this will only occur if a vertex can be encountered via multiple paths traversed in the search. Admittedly, this is rare, but it can occur. It occurs because Blazegraph's predecessor is the first encountered, whereas QLever's is sampled from all possible predecessors in the path.
== Using Multiple Seeds (gas:in Parameters) ==
'''Find all descendants of Richard Burton and Elizabeth Taylor (Q9682) up to 3 levels deep and list their parents'''
This is interesting since Burton and Taylor had children of their own, plus children from other relationships. The seeds/frontiers are multi-generational and multi-marriage on both sides.
Following the path for Burton (3 deep) results in 7 descendants. Following the path for Taylor (3 deep) results in 6 descendants. For both combined, there are 11 descendants.
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?vertex ?depth WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q151973 ; # Richard Burton
gas:in wd:Q34851 ; # Elizabeth Taylor
gas:linkType wdt:P40 ;
gas:maxIterations 3 ;
gas:out ?vertex ;
gas:out1 ?depth }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 11 results <br>
Execution time: 478 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?vertex ?depth ?fromSeed
WHERE {
{ SELECT ?fromSeed ?vertex (MIN(?d) AS ?depth)
WHERE {
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q151973 ;
pathSearch:source wd:Q34851 ;
pathSearch:maxDepth 3 ;
pathSearch:start ?fromSeed ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?fromSeed wdt:P40 ?vertex } }
} # Close the pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
UNION
{ VALUES ?fromSeed { wd:Q151973 wd:Q34851 }
BIND(?fromSeed AS ?vertex)
BIND(0 AS ?d)
} # Close the BIND UNION SERVICE
} GROUP BY ?fromSeed ?vertex # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 13 results <br>
Execution time: 66 msec
=== Why are the number of results different? ===
The result count is different because there are duplicated descendants returned. Richard Burton and Elizabeth Taylor had 2 children from their own marriages - Maria Burton (Q14955495, who was adopted) and Liza Todd (Q14955489). These are repeated - once for each ?fromSeed.
Remember that Blazegraph GAS treats the two seeds (Burton and Taylor) as one merged frontier. A vertex reachable from either seed is visited exactly once, whichever seed's traversal thread gets there first. One row results.
QLever's rewrite does not merge seeds. pathSearches are reported per-seed, and the rewrite adds a ?fromSeed provenance column because it does not silently drop the seed-to-path details.
== Using a Target (gas:target Parameter) ==
'''Walk the ancestral tree of Mary Walker Randolph to show that she is descended from Thomas Jefferson'''
Blazegraph query:
<syntaxhighlight lang="sparql">
PREFIX gas: <http://www.bigdata.com/rdf/gas#>
SELECT ?vertex ?depth ?predecessor WHERE {
SERVICE gas:service {
gas:program gas:gasClass "com.bigdata.rdf.graph.analytics.BFS" ;
gas:in wd:Q65511807 ; # Mary Randolph
gas:target wd:Q11812 ; # Thomas Jefferson
gas:linkType wdt:P40 ;
gas:traversalDirection "Reverse" ;
gas:out ?vertex ;
gas:out1 ?depth }
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 5 results, ending with Thomas Jefferson <br>
Execution time: 292 msec
QLever query:
<syntaxhighlight lang="sparql">
PREFIX pathSearch: <https://qlever.cs.uni-freiburg.de/pathSearch/>
SELECT ?vertex ?depth ?predecessor
WHERE {
{ SELECT ?vertex (MIN(?d) AS ?depth)
WHERE {
{ { BIND(wd:Q65511807 AS ?vertex)
BIND(0 AS ?d)
}
UNION
{ SERVICE pathSearch:
{ _:b0 pathSearch:algorithm pathSearch:allPaths ;
pathSearch:source wd:Q65511807 ;
pathSearch:target wd:Q11812 ;
pathSearch:maxDepth 5 ;
pathSearch:start ?predecessor ;
pathSearch:end ?vertex ;
pathSearch:pathColumn ?pc ;
pathSearch:edgeColumn ?edge
{ SELECT * WHERE { ?vertex wdt:P40 ?predecessor } } # Reverse
} # Close the pathSearch parameter definition
BIND(( ?edge + 1 ) AS ?d)
} # Close the pathSearch SERVICE
} # Close the BIND UNION SERVICE
} GROUP BY ?vertex # Close the embedded SELECT's WHERE
} # Close the embedded SELECT
} ORDER BY ?depth
</syntaxhighlight>
Number of results: 5 results, ending with Thomas Jefferson <br>
Execution time: 185 msec
[[Category:WDQS]]
ek8rw9is9zhsmtuxv6dej0h4o01zqdj
SLO/mediawiki.page change.v1
0
460608
2456780
2455171
2026-09-11T14:46:01Z
APizzata-WMF
47941
added targets and other sections
2456780
wikitext
text/x-wiki
Status: '''[[SLO/Runbook#Drafting|draft]]'''
== Organizational ==
<mark>[[SLO/Template_instructions/Organizational|Instructions]]</mark>
=== Service ===
The [https://stream.wikimedia.org/?doc#/streams/get_v2_stream_mediawiki_page_change_v1 mediawiki.page_change.v1] is a stream produced by EventBus that lands in the [https://datahub.wikimedia.org/dataset/urn:li:dataset:(urn:li:dataPlatform:hive,event.mediawiki_page_change_v1,PROD)/Columns?highlightedPath&is_lineage_mode=false&schemaFilter= event.mediawiki_page_change_v1] Hive table.
The process that populates the [[Kafka]] topic up to [[Data Platform/Systems/Hive|Hive]] is the following:
* [[mw:Extension:EventBus|EventBus]] propagates state changes to the [[EventGate]] instance.
* [[EventGate]] validates the new stream events against a [[Event Platform/Schemas|JSONSchema]], and then produces them to the [[Kafka]] backend.
* [[Data Platform/Systems/Gobblin|Gobblin]] [https://airflow.wikimedia.org/dags/gobblin_event_default/grid?lastrun=running&search=gobblin_event_default job] extracts events from the [[Kafka]] topic and lands them in HDFS.
* Refine job load from HDFS to [[Data Platform/Systems/Hive|Hive]] table.
=== Teams ===
The [[mw:Data_Platform_Engineering/Data_Engineering|Data Engineering]] team is responsible for MediaWiki Content History pipelines and relative tables.
== Architectural ==
=== Environmental dependencies ===
[[File:Page change slo architecture.png|center|frame|'''mediawiki.page_change.v1 architecture and its dependencies''']]
=== Service dependencies ===
For dependencies without an SLO yet, or dependencies that habitually miss their SLO, it's assumed that they maintain their historical performance, or worsen slightly but not dramatically (as recommended by the [[SLO/Template instructions/Service level objectives#Calculate the realistic targets|template instructions]]).
Additionally, all the dependencies levels are split in ''Batch'' and ''Stream'' depending to which kind of process they relate.
==== Hard ====
Degradation of hard dependencies will impact the ability of the system to be reachable or data to be available.
* <u>Stream</u>: Kafka service
* <u>Batch</u>: YARN, Kubernetes (dse-k8s), HDFS
==== Soft ====
Degradation of soft dependencies will impact the system’s ability to present up to date or complete data.
* <u>Stream</u>: MediaWiki (via the EventStreamConfig extension), EventBus ([[SLO/Event Platform|SLO]])
* <u>Batch</u>: all the <u>Stream</u> dependencies, and additionally Gobblin, and Refine ingestion pipeline (no SLO)
==== Indirect ====
Stream
Batch
== Client-facing ==
The table below summarizes the use cases and expectations of the main consumers of this stream.
{| class="wikitable"
|+
!Team
!Use case
!Expectations
|-
|Machine Learning
|The team heavily uses [[changeprop]] to trigger [[Machine Learning/LiftWing|LiftWing]] and compute per-page and per-revision model outputs. Those outputs are often produced to a separate event stream. Downstream uses are:
* [[Search/CirrusSearch|CirrusSearch]] via weighted_tags (article topic, article country, draft topics models)
* Analytics and batch processing for research (revert risk model)
* Product features (the revise tone model is used for precomputed structured tasks)
|Most use cases are not latency sensitive, so having guarantees on completeness is more important than freshness.
|-
|Search
|Search consumes the stream in the [[Search/Update Pipeline|Search Update Pipeline]] (SUP), using a Flink app, to update the search index on revision and re-render changes.
Search offer a 10min [https://grafana.wikimedia.org/d/8xDerelVz/search-update-lag-slo?var-slo_period=7d&orgId=1&from=now-7d&to=now&timezone=browser&var-source=000000017&var-job=search_eqiad&var-threshold=600 SLO on freshness] of the search index relative to the original change in MediaWiki. A slow reconciliation process (with no SLO) also exists for the search index.
|<code>mediawiki.page_change.v1</code> is a critical dependency for SUP, and its freshness should support the existing 10-min SLO. The ideal freshness for <code>mediawiki.page_change.v1</code> would be around 1min.
Very high completeness is not as important, because of the reconciliation mechanism.
|-
|Wikidata Platform
|Wikidata Platform use the stream as a source of notification, to known when to retrieve changes in the Wikidta RDF/triples data model and apply real-time updates to WDQS's database.
Similar to Search, there is a [[SLO/WDQS|10min SLO on freshness for 95% of updates]]. There is no separate reconciliation mechanism, but one is planned for the next iteration of WDQS.
|Freshness expectations are similar to Search. Having ordered events would also be beneficial as an optimization, because of the size of the triplet store.
|}
== Service Level Indicators (SLIs) ==
System health is described by the latency built up in the stream and completeness of the final Hive table:
* '''stream freshness''': latency between the event timestamp and current time of consumption.
* '''batch completeness''': ensures that the page and revision changes available in the [[MariaDB]] are available and accessible.
We measure these two objectives with the following indicators:
* '''Daily freshness SLI for mediawiki.page_change.v1:''' this SLI is identified by the difference between the <code>dt</code> and the Kafka header timestamp for every event available in the [[Kafka]] topic. An event is considered ''fresh'' when the difference between the time timestamps is lower a given threshold.
* '''Monthly completeness SLI for event.mediawiki_page_change_v1:''' this SLI is identified by the completeness of the table. Completeness is a metric defined as the percentage of available <code>revision_id</code> rows for all the Wikis stored in the table. The table is considered complete if the percentage is above a given threshold. Finally, the revisions stored in the table are compared to the ones stored in the <code>wmf_raw.mediawiki_revision</code>.
== Operational ==
<mark>[[SLO/Template_instructions/Operational|Instructions]]</mark>
=== Monitoring ===
==== Batch processing ====
[https://airflow.wikimedia.org/dags/gobblin_event_default/grid Gobblin Airflow hourly airflow task]
[https://airflow.wikimedia.org/dags/refine_to_hive_hourly/grid Refine to Hive hourly airflow task]
==== Stream processing ====
The stream can be monitored on [https://grafana.wikimedia.org/d/000000234/kafka-by-topic?from=now-24h&timezone=utc&to=now&var-datasource=000000006&var-kafka_broker=$__all&var-kafka_cluster=main-eqiad&var-topic=eqiad.mediawiki.page_change.v1&var-topic=codfw.mediawiki.page_change.v1 this] dashboard, to monitor the messages instead it is possible to access to the Kafka ui: [https://kafka.wikimedia.org/ui/clusters/jumbo-eqiad/all-topics/codfw.mediawiki.page_change.v1/messages?limit=100&mode=LATEST codfw] and eqiad.
=== Troubleshooting ===
==== Batch processing ====
Data pipelines have been implemented according to [[Data Platform/Systems/Airflow/Developer guide|Airflow Developer Guide]] guidelines. Generic troubleshooting and operations is described in DPE [[Data Platform Engineering/Ops week|Ops week]] page. Upon failure, data pipelines have to be re-run to backfill data.
==== Stream processing ====
<code>mediawiki.page_change.v1</code> depends on EventBus, Eventgate, and Kafka. Operational errors are expected to be correlated to the performance of either system. The application is deployed on k8s, integrates with the observability platform and follows deployment pipeline operational conventions. More information can be found in the [https://wikitech.wikimedia.org/w/index.php?title=SLO%2FEvent_Platform&wvprov=sticky-header#Troubleshooting Event Platform SLO document].
=== Deployment ===
==== Batch processing ====
Data pipelines DAGs are stored in the airflow-dags repository and are manually deployed on the analytics Airflow instance using scap. The deployment steps are documented on [[Data Platform/Systems/Airflow/Instances|Wikitech]].
==== Stream processing ====
EventStreams are deployed following the [[Deployments on kubernetes#Deploying with helmfile|recommended Kubernetes helmfile deploy pattern]] for services.
==== '''Metric Processing''' ====
To compute the completeness SLO metric an airflow-dag has been deployed, the code it executes is stored in the XXXX repository and is deployed through a CI/CD pipeline after a merge request is merged.
To compute the freshness SLO metric a Flink application has been deployed, the code it executes is stored in the [[gitlab:repos/data-engineering/mediawiki-event-enrichment|mediawiki-event-enrichment]] repository and is deployed through [[Deployments on kubernetes#Deploying with helmfile|helmfile deploy pattern]].
== Service Level Objectives ==
<mark>[[SLO/Template_instructions/Service_level_objectives|Instructions]]</mark>
=== Realistic targets ===
<mark>What are the realistic targets for each SLI? Why?</mark>
==== Streaming ====
95% of events have a latency < 60s.
During testing on different days Obtained the following results:
===== Test 1 2026-09-04 - 2026-09-05 =====
{| class="wikitable sortable"
|+ Event latency distribution (kafka ts − event dt)
! Latency bucket !! Count !! % !! Cumulative %
|-
| 0-5s || 1,536,288 || 95.81% || 95.81%
|-
| 5-10s || 45,131 || 2.81% || 98.62%
|-
| 10-20s || 13,856 || 0.86% || 99.49%
|-
| 20-30s || 3,476 || 0.22% || 99.70%
|-
| 30-40s || 1,524 || 0.10% || 99.80%
|-
| 40-50s || 711 || 0.04% || 99.84%
|-
| 50-60s || 333 || 0.02% || 99.86%
|-
| >60s || 2,206 || 0.14% || 100.00%
|-
! Total !! 1,603,525 !! 100.00% !!
|}
===== Test 2 2026-09-07 - 2026-09-08 =====
{| class="wikitable"
|+ Event latency distribution (kafka_ts − dt)
! Latency bucket !! Count !! % !! Cumulative %
|-
| 0-5s || 1,503,056 || 95.90% || 95.90%
|-
| 5-10s || 42,940 || 2.74% || 98.64%
|-
| 10-20s || 14,177 || 0.90% || 99.55%
|-
| 20-30s || 3,470 || 0.22% || 99.77%
|-
| 30-40s || 1,463 || 0.09% || 99.86%
|-
| 40-50s || 760 || 0.05% || 99.94%
|-
| 50-60s || 412 || 0.03% || 99.94%
|-
| >60s || 985 || 0.06% || 100.00%
|-
! Total !! 1,567,263 !! 100.00% !!
|}
===== Test 3 2026-09-08 - 2026-09-09 =====
{| class="wikitable"
|+ Event latency distribution (kafka_ts − dt)
! Latency bucket !! Count !! % !! Cumulative %
|-
| 0–5s || 1,406,702 || 95.26% || 95.26%
|-
| 5–10s || 43,003 || 2.91% || 98.17%
|-
| 10–20s || 14,145 || 0.96% || 99.13%
|-
| 20–30s || 3,466 || 0.23% || 99.36%
|-
| 30–40s || 1,529 || 0.10% || 99.46%
|-
| 40–50s || 661 || 0.04% || 99.51%
|-
| 50–60s || 313 || 0.02% || 99.53%
|-
| >60s || 6,944 || 0.47% || 100.00%
|-
! Total !! 1,476,763 !! 100.00% !!
|}
===== Test 3 2026-09-09 - 2026-09-10 =====
{| class="wikitable"
|+ Event latency distribution (kafka_ts − dt)
! Latency bucket !! Count !! % !! Cumulative %
|-
| 0–5s || 1,543,493 || 95.92% || 95.92%
|-
| 5–10s || 41,401 || 2.57% || 98.50%
|-
| 10–20s || 13,306 || 0.83% || 99.32%
|-
| 20–30s || 3,375 || 0.21% || 99.53%
|-
| 30–40s || 1,267 || 0.08% || 99.61%
|-
| 40–50s || 425 || 0.03% || 99.64%
|-
| 50–60s || 203 || 0.01% || 99.65%
|-
| >60s || 5,616 || 0.35% || 100.00%
|-
! Total !! 1,609,086 !! 100.00% !!
|}
In the span on 3 months (quarter window), with an average of 1.5 M events a day it would mean:
* setting the target to 99% would mean allowing 1.35 M events with a latency >60s, which is roughly 1 day of data
* setting the target to 97% would mean allowing 4.05 M events with a latency >60s, which is roughly 3 days of data
* setting the target to 95% would mean allowing 6.75 M events with a latency >60s, which is roughly 4,5 days of data
==== Batch ====
99% of <code>event.mediawiki_page_change_v1</code> revision_id, wiki_id combinations with revision_timestamp falling in last month are present in the most current snapshot of <code>wmf_raw.mediawiki_revision</code> table.
=== Ideal targets ===
<mark>What are the ideal targets for each SLI? Why?</mark>
==== Streaming ====
99.9% of events have a latency < 60s: as stated in the [[#Client-facing|Client-facing]] section, both Search and Wikidata value the freshness under 1 minute.
==== Batch ====
99.9% of revisions per wiki_id of last month are present in the most current snapshot of <code>wmf_raw.mediawiki_revision</code> table.
=== Reconciliation ===
<mark>Reconcile the realistic vs. ideal targets, documenting any decisions made along the way.</mark>
<mark>Once the SLO is final, consider collapsing the above three sections.</mark>
<mark>What are the agreed-upon SLOs, for each SLI and each request class?</mark>
<mark>Each SLO should be defined in Sloth; include links here.</mark>
'''Freshness SLO for mediawiki.page_change.v1:''' On at least 95% of the events available in the Kafka topic in a rolling window of 3 months the difference between the <code>dt</code> and the Kafka header timestamp is less or equal than 60 s.
'''Completeness SLO for''' '''event.mediawiki_page_change_v1:''' On at least 90% of months in a rolling window of a year (11 out of 12) the <code>event.mediawiki_page_change_v1</code> table contains more than 99% of the source <code>revision_id</code> rows for all the Wikis, compared to the [[Data Platform/Systems/DB Replica|MariaDB analytics replica]] table <code>wmf_raw.mediawiki_revision</code>, stored in the datalake.
cgvram35t58q3olzsdo0eznd3kf6aom
Talk:SRE/Code Review Culture
1
460617
2456784
2454213
2026-09-11T16:17:25Z
JHathaway (WMF)
28372
/* When do we need code review? */ Reply
2456784
wikitext
text/x-wiki
== When do we need code review? ==
When I started, I was told that every change to puppet required a code review, and that one should never merge one's own puppet CR without a +1 from someone else. So, for instance, currently every change that drains a host from or adds a host to the swift rings (which are very routine changes) gets reviewed by someone else in D-P.
I'm not sure if I was mis-informed, or if culture was/is different in different bits of SRE; is this intended as a change to current practice, or have I misunderstood current practice? [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:27, 2 September 2026 (UTC)
:This is exactly the drift in practices and understanding of the norms we're trying to solve. FWIW, I have always self-merged trivial config changes. For something that can be disruptive and hard to rollback, though, a second pair of eyes are always a good idea - not sure if draining hosts from swift rings falls into that category. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:55, 2 September 2026 (UTC)
:I think the practice of self merging trivial code changes has been inconsistent. I would welcome codifying the practice however, as I think code reviews for trivial or uncontroversial changes and little value. Of course agreeing on how to properly categorize said changes is a bit more difficult. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:17, 11 September 2026 (UTC)
== Pick one reviewer and ping them ==
Perhaps related to what I said above about getting a +1 for everything, we do have a bit of a custom in D-P of asking on our team channel "anyone got time to look at [gerrit link]?", for simple changes that don't need in-depth knowledge (e.g. starting to drain a host from the swift rings); and likewise a culture of helping each other out with reviews if we have some spare attention. This is/seems to be working reasonably well, and I think a "you must pick one person and ping them" system would be an impediment to getting things done. Although, as per previous topic, maybe the answer is "all those routine changes don't need CR any more"
Relatedly, I'd much rather not get separately pinged for a non-urgent CR where someone has tagged me in in Gerrit - gerrit will email me, and that's then on my stack in an async manner, rather than being pinged on IRC (which is more of an interruption - I take an IRC ping of any form as an interrupt, albeit one that can be ignored). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:36, 2 September 2026 (UTC)
:The practice of shopping in team channels for reviews, which every one of us does, is what we're trying to avoid here.
:Noted on your personal preference not to get pinged, but we considered that the gerrit notification mechanism allows too much "not having seen" a patch, and thus delays in the code review process.
:Also, interaction is a positive in code review! We want to encourage direct interaction and less mediated-by-gerrit ones. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:58, 2 September 2026 (UTC)
== Do not comment on a patch without casting a vote ==
I think this is a change to (some) current practice; I get a fair number of code reviews where there are comments that need addressing, and no vote. I understand those to be "please address these comments, and then I will +1". That feels softer/kinder than doing that and adding a -1 vote (and I'm not sure adding a -1 would help much). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:42, 2 September 2026 (UTC)
:Yes, it is a change to some current practice that we consider not optimal. Intentions need to be clear in code review, and also, it's important not to set the norm that a -1 on a patch is a negative judgement. It simply means "needs changes", and we should feel more at ease declaring openly what our intent is in making the review.
:In general, the point is that the reviewer should never leave the author in a limbo. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 14:00, 2 September 2026 (UTC)
:I agree with Matthew, I prefer to comment without -1 as it feels kinder. To me there is no limbo if there is a comment with suggestions. I don't think I've ever left a -2 and use -1 when there is something fundamentally wrong with the change. [[User:AYounsi (WMF)|AYounsi (WMF)]] ([[User talk:AYounsi (WMF)|talk]]) 08:27, 3 September 2026 (UTC)
== Another exception for pushing code without review ==
It's worth to mention and maybe give some examples of "emergency deploys" that obviously doesn't require code reviews: during active serious incidents usually people that are oncall _could_ push changes without reviews (even if it's better if the other oncall person checks it). Not trivial to decide when it's a real emergency situation (but luckily I've never seen this situation abused here). [[User:FFurnari-WMF|FFurnari-WMF]] ([[User talk:FFurnari-WMF|talk]]) 09:00, 2 September 2026 (UTC)
:I had been thinking about this too, but also had decided, well, we generally *do* have multiple people around during big emergencies, and usually fix changes get stamped very quickly since people are paying lots of attention. But it's probably still worth calling out explicitly, but briefly. ✍ [[User:CDanis|CDanis]] 14:28, 2 September 2026 (UTC)
kroenfydc1ejlw7q6600wzgtxcn8tdz
2456786
2456784
2026-09-11T16:38:17Z
JHathaway (WMF)
28372
/* Pick one reviewer and ping them */ Reply
2456786
wikitext
text/x-wiki
== When do we need code review? ==
When I started, I was told that every change to puppet required a code review, and that one should never merge one's own puppet CR without a +1 from someone else. So, for instance, currently every change that drains a host from or adds a host to the swift rings (which are very routine changes) gets reviewed by someone else in D-P.
I'm not sure if I was mis-informed, or if culture was/is different in different bits of SRE; is this intended as a change to current practice, or have I misunderstood current practice? [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:27, 2 September 2026 (UTC)
:This is exactly the drift in practices and understanding of the norms we're trying to solve. FWIW, I have always self-merged trivial config changes. For something that can be disruptive and hard to rollback, though, a second pair of eyes are always a good idea - not sure if draining hosts from swift rings falls into that category. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:55, 2 September 2026 (UTC)
:I think the practice of self merging trivial code changes has been inconsistent. I would welcome codifying the practice however, as I think code reviews for trivial or uncontroversial changes and little value. Of course agreeing on how to properly categorize said changes is a bit more difficult. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:17, 11 September 2026 (UTC)
== Pick one reviewer and ping them ==
Perhaps related to what I said above about getting a +1 for everything, we do have a bit of a custom in D-P of asking on our team channel "anyone got time to look at [gerrit link]?", for simple changes that don't need in-depth knowledge (e.g. starting to drain a host from the swift rings); and likewise a culture of helping each other out with reviews if we have some spare attention. This is/seems to be working reasonably well, and I think a "you must pick one person and ping them" system would be an impediment to getting things done. Although, as per previous topic, maybe the answer is "all those routine changes don't need CR any more"
Relatedly, I'd much rather not get separately pinged for a non-urgent CR where someone has tagged me in in Gerrit - gerrit will email me, and that's then on my stack in an async manner, rather than being pinged on IRC (which is more of an interruption - I take an IRC ping of any form as an interrupt, albeit one that can be ignored). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:36, 2 September 2026 (UTC)
:The practice of shopping in team channels for reviews, which every one of us does, is what we're trying to avoid here.
:Noted on your personal preference not to get pinged, but we considered that the gerrit notification mechanism allows too much "not having seen" a patch, and thus delays in the code review process.
:Also, interaction is a positive in code review! We want to encourage direct interaction and less mediated-by-gerrit ones. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:58, 2 September 2026 (UTC)
:I think Gerrit's email and desktop notifications are sufficient for requesting a review. If a review is not provided or acknowledged in an appropriate amount of time, than synchronous pinging is acceptable. But, it should be avoided as a first action. Pining interrupts a person and is poor fit for colleagues who stretch across time zones. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:38, 11 September 2026 (UTC)
== Do not comment on a patch without casting a vote ==
I think this is a change to (some) current practice; I get a fair number of code reviews where there are comments that need addressing, and no vote. I understand those to be "please address these comments, and then I will +1". That feels softer/kinder than doing that and adding a -1 vote (and I'm not sure adding a -1 would help much). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:42, 2 September 2026 (UTC)
:Yes, it is a change to some current practice that we consider not optimal. Intentions need to be clear in code review, and also, it's important not to set the norm that a -1 on a patch is a negative judgement. It simply means "needs changes", and we should feel more at ease declaring openly what our intent is in making the review.
:In general, the point is that the reviewer should never leave the author in a limbo. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 14:00, 2 September 2026 (UTC)
:I agree with Matthew, I prefer to comment without -1 as it feels kinder. To me there is no limbo if there is a comment with suggestions. I don't think I've ever left a -2 and use -1 when there is something fundamentally wrong with the change. [[User:AYounsi (WMF)|AYounsi (WMF)]] ([[User talk:AYounsi (WMF)|talk]]) 08:27, 3 September 2026 (UTC)
== Another exception for pushing code without review ==
It's worth to mention and maybe give some examples of "emergency deploys" that obviously doesn't require code reviews: during active serious incidents usually people that are oncall _could_ push changes without reviews (even if it's better if the other oncall person checks it). Not trivial to decide when it's a real emergency situation (but luckily I've never seen this situation abused here). [[User:FFurnari-WMF|FFurnari-WMF]] ([[User talk:FFurnari-WMF|talk]]) 09:00, 2 September 2026 (UTC)
:I had been thinking about this too, but also had decided, well, we generally *do* have multiple people around during big emergencies, and usually fix changes get stamped very quickly since people are paying lots of attention. But it's probably still worth calling out explicitly, but briefly. ✍ [[User:CDanis|CDanis]] 14:28, 2 September 2026 (UTC)
kexzwcrobf9s470cdr9hei7rl7z72qp
2456789
2456786
2026-09-11T16:42:44Z
JHathaway (WMF)
28372
/* Guidelines, not rules */ new section
2456789
wikitext
text/x-wiki
== When do we need code review? ==
When I started, I was told that every change to puppet required a code review, and that one should never merge one's own puppet CR without a +1 from someone else. So, for instance, currently every change that drains a host from or adds a host to the swift rings (which are very routine changes) gets reviewed by someone else in D-P.
I'm not sure if I was mis-informed, or if culture was/is different in different bits of SRE; is this intended as a change to current practice, or have I misunderstood current practice? [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:27, 2 September 2026 (UTC)
:This is exactly the drift in practices and understanding of the norms we're trying to solve. FWIW, I have always self-merged trivial config changes. For something that can be disruptive and hard to rollback, though, a second pair of eyes are always a good idea - not sure if draining hosts from swift rings falls into that category. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:55, 2 September 2026 (UTC)
:I think the practice of self merging trivial code changes has been inconsistent. I would welcome codifying the practice however, as I think code reviews for trivial or uncontroversial changes and little value. Of course agreeing on how to properly categorize said changes is a bit more difficult. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:17, 11 September 2026 (UTC)
== Pick one reviewer and ping them ==
Perhaps related to what I said above about getting a +1 for everything, we do have a bit of a custom in D-P of asking on our team channel "anyone got time to look at [gerrit link]?", for simple changes that don't need in-depth knowledge (e.g. starting to drain a host from the swift rings); and likewise a culture of helping each other out with reviews if we have some spare attention. This is/seems to be working reasonably well, and I think a "you must pick one person and ping them" system would be an impediment to getting things done. Although, as per previous topic, maybe the answer is "all those routine changes don't need CR any more"
Relatedly, I'd much rather not get separately pinged for a non-urgent CR where someone has tagged me in in Gerrit - gerrit will email me, and that's then on my stack in an async manner, rather than being pinged on IRC (which is more of an interruption - I take an IRC ping of any form as an interrupt, albeit one that can be ignored). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:36, 2 September 2026 (UTC)
:The practice of shopping in team channels for reviews, which every one of us does, is what we're trying to avoid here.
:Noted on your personal preference not to get pinged, but we considered that the gerrit notification mechanism allows too much "not having seen" a patch, and thus delays in the code review process.
:Also, interaction is a positive in code review! We want to encourage direct interaction and less mediated-by-gerrit ones. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:58, 2 September 2026 (UTC)
:I think Gerrit's email and desktop notifications are sufficient for requesting a review. If a review is not provided or acknowledged in an appropriate amount of time, than synchronous pinging is acceptable. But, it should be avoided as a first action. Pining interrupts a person and is poor fit for colleagues who stretch across time zones. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:38, 11 September 2026 (UTC)
== Do not comment on a patch without casting a vote ==
I think this is a change to (some) current practice; I get a fair number of code reviews where there are comments that need addressing, and no vote. I understand those to be "please address these comments, and then I will +1". That feels softer/kinder than doing that and adding a -1 vote (and I'm not sure adding a -1 would help much). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:42, 2 September 2026 (UTC)
:Yes, it is a change to some current practice that we consider not optimal. Intentions need to be clear in code review, and also, it's important not to set the norm that a -1 on a patch is a negative judgement. It simply means "needs changes", and we should feel more at ease declaring openly what our intent is in making the review.
:In general, the point is that the reviewer should never leave the author in a limbo. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 14:00, 2 September 2026 (UTC)
:I agree with Matthew, I prefer to comment without -1 as it feels kinder. To me there is no limbo if there is a comment with suggestions. I don't think I've ever left a -2 and use -1 when there is something fundamentally wrong with the change. [[User:AYounsi (WMF)|AYounsi (WMF)]] ([[User talk:AYounsi (WMF)|talk]]) 08:27, 3 September 2026 (UTC)
== Another exception for pushing code without review ==
It's worth to mention and maybe give some examples of "emergency deploys" that obviously doesn't require code reviews: during active serious incidents usually people that are oncall _could_ push changes without reviews (even if it's better if the other oncall person checks it). Not trivial to decide when it's a real emergency situation (but luckily I've never seen this situation abused here). [[User:FFurnari-WMF|FFurnari-WMF]] ([[User talk:FFurnari-WMF|talk]]) 09:00, 2 September 2026 (UTC)
:I had been thinking about this too, but also had decided, well, we generally *do* have multiple people around during big emergencies, and usually fix changes get stamped very quickly since people are paying lots of attention. But it's probably still worth calling out explicitly, but briefly. ✍ [[User:CDanis|CDanis]] 14:28, 2 September 2026 (UTC)
== Guidelines, not rules ==
I appreciate that the word "guidelines" is used, but I still think it would be valuable to explicitly note that folks should use their own judgement, and break these rules on occasion. As a an example, if I am reviewing a logic patch, which seems solid, but also includes a small syntax cleanup. I don't think it is valuable to request the patch to be split up, I'll just +1, and perhaps note that next time they should create a separate patch for the syntax piece. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:42, 11 September 2026 (UTC)
iw34xx5vl5iv7p1in1zwhxj1yoebn51
2456792
2456789
2026-09-11T17:03:33Z
JHathaway (WMF)
28372
/* LLM Code reviews */ new section
2456792
wikitext
text/x-wiki
== When do we need code review? ==
When I started, I was told that every change to puppet required a code review, and that one should never merge one's own puppet CR without a +1 from someone else. So, for instance, currently every change that drains a host from or adds a host to the swift rings (which are very routine changes) gets reviewed by someone else in D-P.
I'm not sure if I was mis-informed, or if culture was/is different in different bits of SRE; is this intended as a change to current practice, or have I misunderstood current practice? [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:27, 2 September 2026 (UTC)
:This is exactly the drift in practices and understanding of the norms we're trying to solve. FWIW, I have always self-merged trivial config changes. For something that can be disruptive and hard to rollback, though, a second pair of eyes are always a good idea - not sure if draining hosts from swift rings falls into that category. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:55, 2 September 2026 (UTC)
:I think the practice of self merging trivial code changes has been inconsistent. I would welcome codifying the practice however, as I think code reviews for trivial or uncontroversial changes and little value. Of course agreeing on how to properly categorize said changes is a bit more difficult. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:17, 11 September 2026 (UTC)
== Pick one reviewer and ping them ==
Perhaps related to what I said above about getting a +1 for everything, we do have a bit of a custom in D-P of asking on our team channel "anyone got time to look at [gerrit link]?", for simple changes that don't need in-depth knowledge (e.g. starting to drain a host from the swift rings); and likewise a culture of helping each other out with reviews if we have some spare attention. This is/seems to be working reasonably well, and I think a "you must pick one person and ping them" system would be an impediment to getting things done. Although, as per previous topic, maybe the answer is "all those routine changes don't need CR any more"
Relatedly, I'd much rather not get separately pinged for a non-urgent CR where someone has tagged me in in Gerrit - gerrit will email me, and that's then on my stack in an async manner, rather than being pinged on IRC (which is more of an interruption - I take an IRC ping of any form as an interrupt, albeit one that can be ignored). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:36, 2 September 2026 (UTC)
:The practice of shopping in team channels for reviews, which every one of us does, is what we're trying to avoid here.
:Noted on your personal preference not to get pinged, but we considered that the gerrit notification mechanism allows too much "not having seen" a patch, and thus delays in the code review process.
:Also, interaction is a positive in code review! We want to encourage direct interaction and less mediated-by-gerrit ones. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:58, 2 September 2026 (UTC)
:I think Gerrit's email and desktop notifications are sufficient for requesting a review. If a review is not provided or acknowledged in an appropriate amount of time, than synchronous pinging is acceptable. But, it should be avoided as a first action. Pining interrupts a person and is poor fit for colleagues who stretch across time zones. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:38, 11 September 2026 (UTC)
== Do not comment on a patch without casting a vote ==
I think this is a change to (some) current practice; I get a fair number of code reviews where there are comments that need addressing, and no vote. I understand those to be "please address these comments, and then I will +1". That feels softer/kinder than doing that and adding a -1 vote (and I'm not sure adding a -1 would help much). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:42, 2 September 2026 (UTC)
:Yes, it is a change to some current practice that we consider not optimal. Intentions need to be clear in code review, and also, it's important not to set the norm that a -1 on a patch is a negative judgement. It simply means "needs changes", and we should feel more at ease declaring openly what our intent is in making the review.
:In general, the point is that the reviewer should never leave the author in a limbo. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 14:00, 2 September 2026 (UTC)
:I agree with Matthew, I prefer to comment without -1 as it feels kinder. To me there is no limbo if there is a comment with suggestions. I don't think I've ever left a -2 and use -1 when there is something fundamentally wrong with the change. [[User:AYounsi (WMF)|AYounsi (WMF)]] ([[User talk:AYounsi (WMF)|talk]]) 08:27, 3 September 2026 (UTC)
== Another exception for pushing code without review ==
It's worth to mention and maybe give some examples of "emergency deploys" that obviously doesn't require code reviews: during active serious incidents usually people that are oncall _could_ push changes without reviews (even if it's better if the other oncall person checks it). Not trivial to decide when it's a real emergency situation (but luckily I've never seen this situation abused here). [[User:FFurnari-WMF|FFurnari-WMF]] ([[User talk:FFurnari-WMF|talk]]) 09:00, 2 September 2026 (UTC)
:I had been thinking about this too, but also had decided, well, we generally *do* have multiple people around during big emergencies, and usually fix changes get stamped very quickly since people are paying lots of attention. But it's probably still worth calling out explicitly, but briefly. ✍ [[User:CDanis|CDanis]] 14:28, 2 September 2026 (UTC)
== Guidelines, not rules ==
I appreciate that the word "guidelines" is used, but I still think it would be valuable to explicitly note that folks should use their own judgement, and break these rules on occasion. As a an example, if I am reviewing a logic patch, which seems solid, but also includes a small syntax cleanup. I don't think it is valuable to request the patch to be split up, I'll just +1, and perhaps note that next time they should create a separate patch for the syntax piece. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:42, 11 September 2026 (UTC)
== LLM Code reviews ==
I would appreciate a separate section on LLM code reviews, as I don't think they can be treated the same as code created by a human. Some pieces I would like to see:
An explicit statement on using an LLM, something to the effect of:
<blockquote>
This code was written (mostly|all|some), by an LLM, with this prompt: <PROMPT>. I understand the LLM code and it meets our level of quality.
</blockquote>
Or if that is not the case, be clear about the effort:
<blockquote>
I prompted the LLM with this prompt: <PROMPT>. I haven't carefully reviewed the output.
</blockquote>
Any iterations on the code, should then be done without an LLM, as [https://rfd.shared.oxide.computer/rfd/0576 Oxide suggests].
Empathy and mentoring are a core part of code review, but should largely be absent when reviewing a piece of LLM code. It should be okay to say, "This seems to be some LLM garbage." There should also not be an expectation to provide the same type of feedback, since the LLM is not learning. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 17:03, 11 September 2026 (UTC)
0r3mkrqu6z14wl86gqwl6gvup5hl1st
2456793
2456792
2026-09-11T17:06:33Z
JHathaway (WMF)
28372
/* Do not comment on a patch without casting a vote */ Reply
2456793
wikitext
text/x-wiki
== When do we need code review? ==
When I started, I was told that every change to puppet required a code review, and that one should never merge one's own puppet CR without a +1 from someone else. So, for instance, currently every change that drains a host from or adds a host to the swift rings (which are very routine changes) gets reviewed by someone else in D-P.
I'm not sure if I was mis-informed, or if culture was/is different in different bits of SRE; is this intended as a change to current practice, or have I misunderstood current practice? [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:27, 2 September 2026 (UTC)
:This is exactly the drift in practices and understanding of the norms we're trying to solve. FWIW, I have always self-merged trivial config changes. For something that can be disruptive and hard to rollback, though, a second pair of eyes are always a good idea - not sure if draining hosts from swift rings falls into that category. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:55, 2 September 2026 (UTC)
:I think the practice of self merging trivial code changes has been inconsistent. I would welcome codifying the practice however, as I think code reviews for trivial or uncontroversial changes and little value. Of course agreeing on how to properly categorize said changes is a bit more difficult. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:17, 11 September 2026 (UTC)
== Pick one reviewer and ping them ==
Perhaps related to what I said above about getting a +1 for everything, we do have a bit of a custom in D-P of asking on our team channel "anyone got time to look at [gerrit link]?", for simple changes that don't need in-depth knowledge (e.g. starting to drain a host from the swift rings); and likewise a culture of helping each other out with reviews if we have some spare attention. This is/seems to be working reasonably well, and I think a "you must pick one person and ping them" system would be an impediment to getting things done. Although, as per previous topic, maybe the answer is "all those routine changes don't need CR any more"
Relatedly, I'd much rather not get separately pinged for a non-urgent CR where someone has tagged me in in Gerrit - gerrit will email me, and that's then on my stack in an async manner, rather than being pinged on IRC (which is more of an interruption - I take an IRC ping of any form as an interrupt, albeit one that can be ignored). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:36, 2 September 2026 (UTC)
:The practice of shopping in team channels for reviews, which every one of us does, is what we're trying to avoid here.
:Noted on your personal preference not to get pinged, but we considered that the gerrit notification mechanism allows too much "not having seen" a patch, and thus delays in the code review process.
:Also, interaction is a positive in code review! We want to encourage direct interaction and less mediated-by-gerrit ones. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 13:58, 2 September 2026 (UTC)
:I think Gerrit's email and desktop notifications are sufficient for requesting a review. If a review is not provided or acknowledged in an appropriate amount of time, than synchronous pinging is acceptable. But, it should be avoided as a first action. Pining interrupts a person and is poor fit for colleagues who stretch across time zones. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:38, 11 September 2026 (UTC)
== Do not comment on a patch without casting a vote ==
I think this is a change to (some) current practice; I get a fair number of code reviews where there are comments that need addressing, and no vote. I understand those to be "please address these comments, and then I will +1". That feels softer/kinder than doing that and adding a -1 vote (and I'm not sure adding a -1 would help much). [[User:MVernon (WMF)|MVernon (WMF)]] ([[User talk:MVernon (WMF)|talk]]) 08:42, 2 September 2026 (UTC)
:Yes, it is a change to some current practice that we consider not optimal. Intentions need to be clear in code review, and also, it's important not to set the norm that a -1 on a patch is a negative judgement. It simply means "needs changes", and we should feel more at ease declaring openly what our intent is in making the review.
:In general, the point is that the reviewer should never leave the author in a limbo. [[User:GLavagetto (WMF)|GLavagetto (WMF)]] ([[User talk:GLavagetto (WMF)|talk]]) 14:00, 2 September 2026 (UTC)
:I agree with Matthew, I prefer to comment without -1 as it feels kinder. To me there is no limbo if there is a comment with suggestions. I don't think I've ever left a -2 and use -1 when there is something fundamentally wrong with the change. [[User:AYounsi (WMF)|AYounsi (WMF)]] ([[User talk:AYounsi (WMF)|talk]]) 08:27, 3 September 2026 (UTC)
:When I am trying to understand a piece of code, I often ask questions. I think a neutral 0 is the most appropriate for a question. Also, if we are going to be more explicit about labeling, shouldn't we align with [https://www.mediawiki.org/wiki/Gerrit/Code_review#Complete_the_review mediawiki's usage]? [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 17:06, 11 September 2026 (UTC)
== Another exception for pushing code without review ==
It's worth to mention and maybe give some examples of "emergency deploys" that obviously doesn't require code reviews: during active serious incidents usually people that are oncall _could_ push changes without reviews (even if it's better if the other oncall person checks it). Not trivial to decide when it's a real emergency situation (but luckily I've never seen this situation abused here). [[User:FFurnari-WMF|FFurnari-WMF]] ([[User talk:FFurnari-WMF|talk]]) 09:00, 2 September 2026 (UTC)
:I had been thinking about this too, but also had decided, well, we generally *do* have multiple people around during big emergencies, and usually fix changes get stamped very quickly since people are paying lots of attention. But it's probably still worth calling out explicitly, but briefly. ✍ [[User:CDanis|CDanis]] 14:28, 2 September 2026 (UTC)
== Guidelines, not rules ==
I appreciate that the word "guidelines" is used, but I still think it would be valuable to explicitly note that folks should use their own judgement, and break these rules on occasion. As a an example, if I am reviewing a logic patch, which seems solid, but also includes a small syntax cleanup. I don't think it is valuable to request the patch to be split up, I'll just +1, and perhaps note that next time they should create a separate patch for the syntax piece. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 16:42, 11 September 2026 (UTC)
== LLM Code reviews ==
I would appreciate a separate section on LLM code reviews, as I don't think they can be treated the same as code created by a human. Some pieces I would like to see:
An explicit statement on using an LLM, something to the effect of:
<blockquote>
This code was written (mostly|all|some), by an LLM, with this prompt: <PROMPT>. I understand the LLM code and it meets our level of quality.
</blockquote>
Or if that is not the case, be clear about the effort:
<blockquote>
I prompted the LLM with this prompt: <PROMPT>. I haven't carefully reviewed the output.
</blockquote>
Any iterations on the code, should then be done without an LLM, as [https://rfd.shared.oxide.computer/rfd/0576 Oxide suggests].
Empathy and mentoring are a core part of code review, but should largely be absent when reviewing a piece of LLM code. It should be okay to say, "This seems to be some LLM garbage." There should also not be an expectation to provide the same type of feedback, since the LLM is not learning. [[User:JHathaway (WMF)|JHathaway (WMF)]] ([[User talk:JHathaway (WMF)|talk]]) 17:03, 11 September 2026 (UTC)
m0f688bsvekdododp5egnmcsct8exvc
News/2026 Commons links tables database split
0
460641
2456794
2455713
2026-09-11T17:06:59Z
BryanDavis
1604
/* Why are we doing this? */ add some phab links
2456794
wikitext
text/x-wiki
{{Data Services nav}}
This page contains information about the '''split of Commons links tables to the new <code>x4</code> database cluster'''.
== What is changing? ==
Tables holding Commons links information are being split to a new separate [[MariaDB#Categories_and_clusters|database cluster]]. Tools querying those tables via the [[Help:Wiki Replicas|wiki replicas]] must be adapted to connect to the new cluster instead.
The following tables are affected by this move:
* linktarget
* externallinks
* pagelinks
* templatelinks
* categorylinks
* collation
* imagelinks
* globalimagelinks
* iwlinks
* existencelinks
* langlinks
In addition, the <code>page</code> and <code>redirect</code> tables will continue to exist on both the existing core (<code>s4</code>) and extension (<code>x4</code>) clusters.
== Timeline ==
* {{Done}} 2026-09-08: New hostnames announced and ready for use.
* {{Done}} 2026-09-08: The <code>x4</code> section is split and MediaWiki is changed to write there. The old tables on the <code>s4</code> cluster will no longer be updated and will eventually be dropped.
* TBC: <code>x4</code> is set up on the Wiki Replicas as a separate section. Queries will start returning up-to-date information again, but no longer have access to <code>s4</code> tables.
* TBC: The old tables in the <code>s4</code> cluster will be dropped from the Wiki Replicas.
== What should I do? ==
If you maintain a tool that queries the Commons links tables listed above, you need to update your code to connect to a separate database cluster for those:
{| class="wikitable"
|+
!Old hostname
!New hostname
|-
|<code>commonswiki.analytics.db.svc.wikimedia.cloud</code>
|<code>links.commonswiki.analytics.db.svc.wikimedia.cloud</code>
|-
|<code>commonswiki.web.db.svc.wikimedia.cloud</code>
|<code>links.commonswiki.web.db.svc.wikimedia.cloud</code>
|-
|<code>testcommonswiki.analytics.db.svc.wikimedia.cloud</code>
|<code>links.testcommonswiki.analytics.db.svc.wikimedia.cloud</code>
|-
|<code>testcommonswiki.web.db.svc.wikimedia.cloud</code>
|<code>links.testcommonswiki.web.db.svc.wikimedia.cloud</code>
|}
In addition, any code performing JOINs between those tables and other <code>commonswiki</code> (or <code>testcommonswiki</code>) will need to be changed to perform the same operations in code instead.
== Why are we doing this? ==
{{Tracked|T343131}}{{Tracked|T398709}}
The Commons database has been growing too fast, and so this change is necessary to allow growth and to improve performance.
== See also ==
* [[Help:Wiki Replicas]]
* [[listarchive:list/cloud-announce@lists.wikimedia.org/thread/E4RXH3F3WTBWSGZDN6BZPJGEUY2IM4KP/|[Cloud-announce] Upcoming: Commons links tables move to dedicated cluster (x4)]]
{{Help:Cloud Services communication}}
3hxybe7rgd91m0enplz0k7htp2toanf
Deployments/Archive/2026/09
0
460652
2456812
2026-09-12T02:00:27Z
DeploymentCalendarTool
20896
Add last week
2456812
wikitext
text/x-wiki
==Week of September 07==
==={{Deployment_day|date=2026-09-06}}===
{{Deployment calendar event card
|when=2026-09-06 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
==={{Deployment_day|date=2026-09-07}}===
{{Deployment calendar event card
|when=2026-09-07 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|abijeet|abijeet}}
{{deploy|type=config|gerrit=1329314|title=ArticleGuidance: Configure feedback links to local talk pages|status=}} - {{phabricator|T433483}}
{{ircnick|Hamishcz|Hamish}}
{{deploy|type=config|gerrit=1335706|title=thwikibooks: update tagline and wordmark|status=}} - {{phabricator|T436426}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-07 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-07 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|Tran|Tran}}
{{deploy|type=1.47.0-wmf.18|gerrit=1337450|title=SI: Use new InfoChip text class in case status updater|status=}} - {{phabricator|T437020}}
{{ircnick|Krinkle|Krinkle}}
{{deploy|type=1.47.0-wmf.18|gerrit=1335752|title=Fix empty bases in m(under{{!}}over)|status=}} - {{phabricator|T436876}}
{{deploy|type=1.47.0-wmf.18|gerrit=1336611|title=Use Core-compatible output for cancellation|status=}} - {{phabricator|T379359}}
{{deploy|type=1.47.0-wmf.18|gerrit=1336609|title=Page: Fix absence caching when combined with multiple properties|status=}} - {{phabricator|T297300}}
{{deploy|type=1.47.0-wmf.18|gerrit=1336613|title=User: Use ActorStore on User::load|status=}} - {{phabricator|T428517}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-07 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-07 08:30 SF
|length=0.5
|window=Wikimedia Portals Update
|who={{ircnick|jan_drewniak|Jan Drewniak}}
|what=Weekly window for the portals page: https://www.wikipedia.org/
}}
{{Deployment calendar event card
|when=2026-09-07 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-07 10:00 SF
|length=0.5
|window=Wikidata Query Service weekly deploy
|who={{ircnick|ryankemper|Ryan}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-07 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|Daimona|Daimona}}
{{deploy|type=config|gerrit=1326400|title=Stop setting wgCampaignEventsEnableWorklists|status=}} - {{phabricator|T429510}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-07 14:00 SF
|length=2
|window=Weekly Security deployment window
|who={{ircnick|alexsanford|Alex}}, {{ircnick|Reedy|Sam}}, {{ircnick|sbassett|Scott}}, {{ircnick|Maryum|Maryum}}, {{ircnick|manfredi|Manfredi}}
|what=Held deployment window for Security-team related deploys.
}}
{{Deployment calendar event card
|when=2026-09-07 16:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-07 19:00 SF
|length=1
|window=Automatic branching of MediaWiki, extensions, skins, and vendor – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Branch <code>wmf/1.47.0-wmf.19</code>
}}
{{Deployment calendar event card
|when=2026-09-07 19:00 SF
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-07 20:00 SF
|length=1
|window=Automatic deployment of MediaWiki, extensions, skins, and vendor to testwikis only – see [[Heterogeneous deployment/Train deploys]]
|who=N/A
|what=Deploy <code>wmf/1.47.0-wmf.19</code> to testwikis
}}
{{Deployment calendar event card
|when=2026-09-07 21:00 SF
|length=1
|window=Automatic removal of all obsolete MediaWiki versions from the deployment and bare metal servers (except the most-recent obsolete version)
|who=N/A
|what=Runs <code>scap clean auto</code>
}}
{{Deployment calendar event card
|when=2026-09-07 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-07 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-08}}===
{{Deployment calendar event card
|when=2026-09-08 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|hamishcz|Hamish}}
{{deploy|type=config|gerrit=1335706|title=thwikibooks: update tagline and wordmark|status=}} - {{phabricator|T436426}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-08 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-08 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-08 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|Tran|Tran}}
{{deploy|type=1.47.0-wmf.18|gerrit=1337610|title=Split out edit and block-based filters from activity filters|status=}} - {{phabricator|T436508}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-08 07:00 SF
|length=0.5
|window=Test Kitchen UI Deployment Window
|who=Experimentation Platform Team
|what=Deployment of Test Kitchen UI (fka MPIC)
}}
{{Deployment calendar event card
|when=2026-09-08 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-08 08:00 SF
|length=1
|window=SRE Collaboration Services office hours
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=Services including Gerrit, Phorge (Phabricator), GitLab
}}
{{Deployment calendar event card
|when=2026-09-08 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-08 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-08 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|dduvall|Dan}}, {{ircnick|dancy|Ahmon}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.18->1.47.0-wmf.19|1.47.0-wmf.18|1.47.0-wmf.18}}
* group0 to [[mw:MediaWiki_1.47/wmf.19|1.47.0-wmf.19]]
* '''Blockers: {{phabricator|T430838}}'''
}}
{{Deployment calendar event card
|when=2026-09-08 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|AaronSchulz|AaronSchulz}}
{{deploy|type=config|gerrit=1326425|title=Add wmf-analytics-commons external module to commonswiki|status=}} - {{phabricator|T434927}}
{{ircnick|sbassett, aranyap|sbassett, aranyap}}
{{deploy|type=1.47.0-wmf.19|gerrit=1337997|title=Filter out non-http(s) license urls|status=}} - {{phabricator|T435999}}
{{deploy|type=1.47.0-wmf.19|gerrit=1337998|title=Filter out non-http(s) license urls|status=}} - {{phabricator|T435999}}
{{deploy|type=1.47.0-wmf.19|gerrit=1337996|title=Filter out non-http(s) license urls|status=}} - {{phabricator|T435999}}
{{ircnick|hamishcz|Hamish}}
{{deploy|type=config|gerrit=1335706|title=thwikibooks: update tagline and wordmark|status=}} - {{phabricator|T436426}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-08 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-08 19:00 SF
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-08 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-09}}===
{{Deployment calendar event card
|when=2026-09-09 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|hamishcz|Hamish}}
{{deploy|type=config|gerrit=1335706|title=thwikibooks: update tagline and wordmark|status=}} - {{phabricator|T436426}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-09 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-09 04:00 SF
|length=1
|window=[[mw:Services|Services]] – [[Citoid]] / [[Zotero]]
|who=Marielle ({{ircnick|mvolz}})
|what=See [[mw:Citoid|Citoid]]
}}
{{Deployment calendar event card
|when=2026-09-09 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|stephanebisson|Stephane Bisson}}
{{deploy|type=config|gerrit=1334945|title=ArticleGuidance: Add the redirect configuration keys|status=}} - {{phabricator|T434487}}
{{deploy|type=1.47.0-wmf.19|gerrit=1337983|title=Replace experiment with instrument and config-driven redirect|status=}} - {{phabricator|T434487}}
{{ircnick|Hide_on_rosie|Hide_on_rosie}}
{{deploy|type=config|gerrit=1338151|title=core-Namespaces.php: Disallow indexing on talk namespaces on ukwiki|status=}} - {{phabricator|T437409}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-09 07:00 SF
|length=1
|window=Wikifunctions Services UTC Afternoon
|who=Abstract Wikipedia team (Africa, Europe, Eastern Americas)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-09 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-09 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-09 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|dduvall|Dan}}, {{ircnick|dancy|Ahmon}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.19|1.47.0-wmf.18->1.47.0-wmf.19|1.47.0-wmf.18}}
* group1 to [[mw:MediaWiki_1.47/wmf.19|1.47.0-wmf.19]]
* '''Blockers: {{phabricator|T430838}}'''
}}
{{Deployment calendar event card
|when=2026-09-09 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|sbassett, aranyap|sbassett, aranyap}}
{{deploy|type=1.47.0-wmf.18|gerrit=1338278|title=Revert^2 "Filter out non-http(s) license urls"|status=}}
{{deploy|type=1.47.0-wmf.19|gerrit=1338279|title=Revert^2 "Filter out non-http(s) license urls"|status=}}
{{deploy|type=1.47.0-wmf.18|gerrit=1338280|title=Revert^2 "Filter out non-http(s) license urls"|status=}}
{{deploy|type=1.47.0-wmf.19|gerrit=1338281|title=Revert^2 "Filter out non-http(s) license urls"|status=}}
{{deploy|type=1.47.0-wmf.18|gerrit=1338282|title=Revert^2 "Filter out non-http(s) license urls"|status=}}
{{deploy|type=1.47.0-wmf.19|gerrit=1338283|title=Revert^2 "Filter out non-http(s) license urls"|status=}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-09 14:00 SF
|length=1
|window=Wikifunctions Services UTC Late
|who=Abstract Wikipedia team (North and South America)
|what=Wikifunctions back-end k8s services
}}
{{Deployment calendar event card
|when=2026-09-09 15:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-09 19:00 SF
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-09 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-09 23:00 SF
|length=0.5
|window=Primary database switchover
|who={{ircnick|marostegui|Manuel Arostegui}}, {{ircnick|cezmunsta|Ceri Williams}}, {{ircnick|federico3|Federico Ceratto}}
|what=Held deployment window for database primary masters maintenance
}}
==={{Deployment_day|date=2026-09-10}}===
{{Deployment calendar event card
|when=2026-09-10 00:00 SF
|length=1
|window=[[Backport windows|UTC morning backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Amir1|Amir}}, {{ircnick|urbanecm|Martin}}, {{ircnick|awight|Adam}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-10 03:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC mid-day)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-10 05:00 SF
|length=1
|window=Mobileapps/RESTBase/Wikifeeds
|who=Content Transform Team
|what=Content transform team node services (mobileapps/wikifeeds)
}}
{{Deployment calendar event card
|when=2026-09-10 06:00 SF
|length=1
|window=[[Backport windows|UTC afternoon backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|Lucas_WMDE|Lucas}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-10 07:30 SF
|length=0.5
|window=Test Kitchen Experiment Deployment Window
|who=Test Kitchen
|what=Automatic start/stop of active experiments and instruments managed by [[Test Kitchen]].
}}
{{Deployment calendar event card
|when=2026-09-10 09:00 SF
|length=1
|window=[[Puppet request window]]<br/><small>'''(Max 6 patches)'''</small>
|who={{ircnick|jhathaway|JHathaway}}, {{ircnick|rzl|Reuven}}
|what={{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to Puppet change''
}}
{{Deployment calendar event card
|when=2026-09-10 10:00 SF
|length=1
|window=Cloud Services/Technical Documentation weekly deploy (Toolhub, Developer portal, Striker)
|who={{ircnick|bd808}}
|what=...
}}
{{Deployment calendar event card
|when=2026-09-10 10:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC late)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
{{Deployment calendar event card
|when=2026-09-10 11:00 SF
|length=2
|window=MediaWiki train - Utc-7 Version
|who={{ircnick|dduvall|Dan}}, {{ircnick|dancy|Ahmon}}
|what=[[mw:MediaWiki 1.47/Roadmap#Schedule for the deployments|1.47 schedule]]
{{DeployOneWeekMini|1.47.0-wmf.19|1.47.0-wmf.19|1.47.0-wmf.18->1.47.0-wmf.19}}
* group2 to [[mw:MediaWiki_1.47/wmf.19|1.47.0-wmf.19]]
* '''Blockers: {{phabricator|T430838}}'''
}}
{{Deployment calendar event card
|when=2026-09-10 13:00 SF
|length=1
|window=[[Backport windows|UTC late backport window]]<br/><small>'''Your patch may or may not be deployed at the sole discretion of the deployer'''</small>
|who={{ircnick|RoanKattouw|Roan}}, {{ircnick|urbanecm|Martin}}, {{ircnick|TheresNoTime|Sammy}}, {{ircnick|kindrobot|Stef}}, {{ircnick|cjming|Clare}}
|what={{ircnick|jan_drewniak|Jan Drewniak}}
* [Portal] {{gerrit|1338975}} Portal banner deploy
{{ircnick|obvious-dust|obvious-dust}}
{{deploy|type=config|gerrit=1328215|title=Remove mode from RestModuleOverrides|status=}} - {{phabricator|T434267}}
{{deploy|type=config|gerrit=1328215|title=Remove mode from RestModuleOverrides|status=}} - {{phabricator|T434267}}
{{ircnick|tgr|Gergő}}
{{deploy|type=config|gerrit=1330446|title=CommonSettings: Use a restrictive, eval-free CSP for auth.wikimedia.org|status=not done}} - {{phabricator|T419684}}
{{ircnick|irc-nickname|Requesting Developer}}
* ''Gerrit link to backport or config change''
}}
{{Deployment calendar event card
|when=2026-09-10 14:00 SF
|length=1
|window=Readers deployment window
|who=Readers
|what=NOTE: often skipped, the reader teams do not typically check IRC so assume this is not being used if 5 minutes past the start
}}
{{Deployment calendar event card
|when=2026-09-10 19:00 SF
|length=1
|window=Automatic deployment of MediaWiki to pretrain wikis - see [[mw:Pretrain]]
|who=N/A
|what=Deploy <code>wmf/next</code> to pretrain environment
}}
{{Deployment calendar event card
|when=2026-09-10 23:00 SF
|length=1
|window=[[MediaWiki_On_Kubernetes#How_to_manage_changes_to_the_infrastructure|MediaWiki infrastructure]] (UTC early)
|who=SRE team
|what=MediaWiki-related infrastructure changes that need a kubernetes deployment.
}}
==={{Deployment_day|date=2026-09-11}}===
{{Deployment calendar event card
|when=2026-09-11 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
{{Deployment calendar event card
|when=2026-09-11 04:00 SF
|length=0.5
|window=GitLab version upgrades
|who={{ircnick|jelto|Jelto}}, {{ircnick|arnoldokoth|Arnold}}, {{ircnick|mutante|Daniel}}, {{ircnick|arnaudb|Arnaud}}
|what=GitLab version upgrades
}}
==={{Deployment_day|date=2026-09-12}}===
{{Deployment calendar event card
|when=2026-09-12 00:00 SF
|length=24
|window=No deploys all day! See [[Deployments/Emergencies]] if things are broken.
|who=
|what=No Deploys
}}
5hmxm4g0fb3a5ygk7lzhhhzn943dw1i
Help:Toolforge/Push to deploy
12
460653
2456816
2026-09-12T08:55:53Z
Taavi
13997
search term
2456816
wikitext
text/x-wiki
#REDIRECT [[Help:Toolforge/Deploy your tool]]
c83ov827v2lnlophaiw5g4qsmks52xh
Push to deploy
0
460654
2456817
2026-09-12T08:56:07Z
Taavi
13997
search term
2456817
wikitext
text/x-wiki
#REDIRECT [[Help:Toolforge/Deploy your tool]]
c83ov827v2lnlophaiw5g4qsmks52xh